← MDifying document to Markdown converter
Papers are the hardest documents on this site to convert well, and worth being honest about. A modern preprint with a real text layer converts cleanly. A two-column journal PDF from 2003, scanned from paper, does not — and knowing which one you have before you paste it into a model saves an hour of confusion.
Convert a file now — it is free and runs in your browser
Try selecting a sentence in your PDF reader. If the text highlights, it converts directly. If nothing selects, the page is a picture of paper and you need the OCR route below.
Section headings should come out as headings and tables as tables. Running headers and page numbers are removed automatically, so the journal name does not repeat forty times through the output.
When a PDF has no text to extract, MDifying offers to read it off the page images instead. That runs on your machine too. It recognises English and Spanish, it is slower, and it is approximate — read the result against the original before citing anything from it.
References, acknowledgements and author affiliations are usually noise for whatever you are about to ask. Cutting them is a few seconds in a text editor and buys back context window.
Two-column layouts are the honest weak point. Lines are grouped by their position down the page, so a two-column paper can interleave the columns into text that reads like nonsense. If that happens, the result is obviously wrong rather than subtly wrong — which is at least a failure you can see. Equations and figures are not reconstructed either; the words around them are.
Unpublished work, material under embargo, and anything covered by a data agreement. Because the conversion runs in your browser, the file never reaches a server — ours or anyone's — which is a materially different position from a converter that uploads and deletes after thirty days.