PDF → Markdown
Converting a novel from PDF (.pdf) to Markdown (.md) is a routine step in practice — Markdown is plain text, so any editor opens it and it versions cleanly. Going from PDF to Markdown does raise extra questions worth thinking about: you lose metadata such as title, author and language; exact pagination and element positions. Below we explain in detail how PDF becomes Markdown and what to watch out for, and the page ends with a converter that does the job right in your browser — nothing is uploaded, the file never leaves your machine.
What is PDF?
PDF records where things are drawn on a page rather than a paragraph structure, and that is what lets it freeze a layout: fonts, columns and figure positions look identical on every device and print exactly as they appear on screen. The trade-off is that it cannot reflow and the reader cannot change the font size, so long text means zooming and panning on a small screen. It suits work whose layout must be preserved — papers, contracts, scans, print-ready files — and documents that need to be archived unchanged. Read the PDF format guide
| Extensions | .pdf |
|---|---|
| Typical use | Documents, papers, contracts and scans where the layout must stay exact; also a print and archival format. |
What is Markdown?
Markdown is plain text plus markers (# headings, ** bold, | tables) — effectively a superset of .txt. Any text editor opens it, and it versions cleanly in Git. Its expressiveness is deliberately limited: no pagination, no headers or footers, no precise control over type size or position, and no complex layouts. Those trade-offs make Markdown an excellent authoring and conversion source, but as a final reading format its experience depends entirely on which app opens it. Read the Markdown format guide
| Extensions | .md |
|---|---|
| Typical use | Authoring, notes and technical docs; a clean intermediate format before converting to EPUB or TXT. |
How to convert PDF to Markdown
Line breaks in PDF are mostly visual, not paragraph ends: one paragraph may arrive as five separate text blocks, each with its own hard return. Merging them back requires heuristics based on line width and sentence endings. A scanned PDF has no text layer at all, so extraction returns nothing until you run OCR.
Markdown output keeps text and only the most basic hierarchy (# headings, paragraphs, simple lists). Pagination, headers and footers, comments, tracked changes, precise type sizes and text wrapping are all gone — it is a content archive, not a layout archive.
Layout must be rebuilt: the paragraph, column and figure positions in the source were arranged for a fixed page and mean nothing once the target reflows. The converter can only re-stitch a text stream, so paragraphs may merge wrongly, columns may interleave and figures may land mid-sentence. Read the first few chapters afterwards, watching dialogue breaks and figure placement.
The source is one full-page image per page while the target reflows: each image is inserted in sequence, and with no text layer you end up with an image stream. The result is usually huge and unsearchable — a worse reading experience than the original.
Metadata in the PDF source — title, author, publisher, language, table of contents, cover — has nowhere to live in the target, so converting to Markdown loses all of it.
Convert PDF to Markdown
PDF → MARKDOWN
The conversion runs entirely in your browser — your file is never uploaded.
A novel reader that syncs your progress across every platform
Tide Reader opens TXT and EPUB, detects chapters on import, lays the text out the way you like it, and syncs your reading position across Windows, macOS, iOS and Android. Drop the converted file in and start reading.
Download Tide ReaderTXT and EPUB; local-first reading, with optional cloud sync and WebDAV.
