Skip to main content
All articles
Format conversion

PDF → MOBI/AZW3

Converting a novel from PDF to MOBI/AZW3: what each format is, what gets lost, the pitfalls for long fiction, and the right tools.

Tide Reader Editorial6 min read

What is PDF?

PDF is Adobe's fixed-layout document format. It records what to draw at which coordinate on a page rather than which paragraph a sentence belongs to. The layout is therefore identical everywhere, which suits contracts, papers and scans. The cost is that PDF is inherently hostile to reflow: line breaks are usually visual, extracted text often carries hard returns and scrambled column order, and scanned PDFs contain images only, with no text layer at all.

Extensions.pdf
Typical useDocuments, papers, contracts and scans where the layout must stay exact; also a print and archival format.

What is MOBI/AZW3?

Mobipocket (.mobi) was long the default format Amazon accepted for Kindle delivery. AZW3 (also called KF8) is its successor with fuller HTML5/CSS support, and .azw is an earlier Amazon container. All three reflow and carry navigation and metadata, so they behave as one family. They are proprietary to Amazon's ecosystem, and since 2022 Amazon sends KFX/EPUB instead, so .mobi is mostly a legacy format now.

Extensions.mobi.azw3.azw
Typical useReading on Kindle devices and the Kindle app; one of the most common proprietary formats in existing libraries.

What to watch out for when converting PDF to MOBI/AZW3

Extracting from the source

Line breaks in PDF are mostly visual, not paragraph ends: one paragraph may arrive as five separate text blocks, each with its own hard return. Merging them back requires heuristics based on line width and sentence endings. A scanned PDF has no text layer at all, so extraction returns nothing until you run OCR.

Writing the target format

Producing .mobi/.azw3 requires Amazon's toolchain or Calibre. AZW3 supports reasonable CSS while legacy .mobi supports very little; for newer devices, target AZW3 — or simply deliver an EPUB.

Where the two formats interact

  • Layout must be rebuilt: the paragraph, column and figure positions in the source were arranged for a fixed page and mean nothing once the target reflows. The converter can only re-stitch a text stream, so paragraphs may merge wrongly, columns may interleave and figures may land mid-sentence. Read the first few chapters afterwards, watching dialogue breaks and figure placement.
  • The source is one full-page image per page while the target reflows: each image is inserted in sequence, and with no text layer you end up with an image stream. The result is usually huge and unsearchable — a worse reading experience than the original.
  • The source has no chapter concept (comics and PDFs are page-based) while the target expects navigation. The converter can only split mechanically by page or fixed length, producing a semantically empty TOC. Ignore it if you do not need navigation, but do not expect real chapters to appear.

What gets lost

  • Exact pagination and element positions

What you gain

  • Reflow: text re-lays-out for any screen and font size
  • A real table of contents with chapter jumps
  • Full support in mainstream readers and libraries

Post-conversion checklist

  1. Convert the first three chapters as a trial and check for lost text, interleaved lines or misplaced blocks before doing the whole book.
  2. Check chapter splitting: the TOC entry count should match the real chapter count.
  3. Jump to five random spots and search a character name to confirm the text is fully searchable.
  4. Confirm illustration count and placement — check that images are not all dumped at the end of a chapter.
  5. Fill in title, author and language — a wrong language tag breaks typesetting such as Chinese line breaking.

Converting novels: what is different

Fiction is the most common and the most peculiar kind of e-book: hundreds or thousands of chapters, almost no figures, often nothing but a TXT source, and Chinese sources that frequently carry anti-piracy garbling. These are the points that bite hardest.

  • A long novel can run to hundreds or thousands of chapters, so the table of contents decides whether the result is usable. Normalise chapter headings to one recognisable pattern first — say "Chapter N Title" alone on its own line — or the target will miss chapters, merge them, or collapse the whole book into one.
  • Volume markers ("Volume One", "Volume Two") usually degrade to ordinary headings, losing the volume/chapter hierarchy. If that matters, promote volume headings one level in an editor afterwards so they are not siblings of chapter titles.

Tools for converting PDF to MOBI/AZW3

Any of the tools below can handle this direction; pick one by your operating system and how many files you have. Read the previous section first — doing the steps out of order (converting a scanned file before running OCR, for instance) usually means starting over.

ocrmypdf / Tesseract

FreeCommand line

An open-source way to add a text layer (OCR) to scanned PDFs, with Chinese support. Scanned files must go through this before any text-extraction conversion, otherwise the output is blank. Expect to proofread the recognised characters.

Platforms: Windows / macOS / Linux (command line)ocrmypdf.readthedocs.io ↗

Stirling PDF

FreeOnline service

An open-source PDF toolbox that splits, merges, rotates, compresses and extracts text in the browser, and can be self-hosted so files never leave your network. Good for the PDF side; for reflowable output Calibre remains the better choice.

Platforms: Browser (self-hostable)stirlingpdf.io ↗

Kindle Previewer

FreeDesktop app

Amazon's official tool: it converts EPUB/DOCX to AZW3 and previews the result exactly as a real Kindle would render it. It is the most reliable way to verify how a conversion actually looks on a Kindle, and avoids hand-building the proprietary container.

Platforms: Windows / macOSkdp.amazon.com/en_US/help/topic/G202131100 ↗

Calibre

FreeDesktop app

The de-facto standard for e-book conversion, free and open source. It converts EPUB, MOBI, AZW3, PDF, TXT, FB2, DOCX, HTML and nearly everything else, generating tables of contents, fetching metadata and batch-processing whole libraries. The interface is utilitarian, PDF output is mediocre, and the option set takes time to learn.

Platforms: Windows / macOS / Linuxcalibre-ebook.com ↗

Common questions

What is lost when converting PDF to MOBI/AZW3?

Exact pagination and element positions.

Which tool should I use to convert PDF to MOBI/AZW3?

Recommended: ocrmypdf / Tesseract, Stirling PDF, Kindle Previewer, Calibre. ocrmypdf / Tesseract fits this pair best. An open-source way to add a text layer (OCR) to scanned PDFs, with Chinese support. Scanned files must go through this before any text-extraction conversion, otherwise the output is blank. Expect to proofread the recognised characters.

What should I watch out for when converting PDF to MOBI/AZW3?

Layout must be rebuilt: the paragraph, column and figure positions in the source were arranged for a fixed page and mean nothing once the target reflows. The converter can only re-stitch a text stream, so paragraphs may merge wrongly, columns may interleave and figures may land mid-sentence. Read the first few chapters afterwards, watching dialogue breaks and figure placement. The source is one full-page image per page while the target reflows: each image is inserted in sequence, and with no text layer you end up with an image stream. The result is usually huge and unsearchable — a worse reading experience than the original. The source has no chapter concept (comics and PDFs are page-based) while the target expects navigation. The converter can only split mechanically by page or fixed length, producing a semantically empty TOC. Ignore it if you do not need navigation, but do not expect real chapters to appear.

Converted it — now what do you read it with?

Tide Reader opens TXT and EPUB, detects chapters on import, lays the text out the way you like it, and syncs your reading position across Windows, macOS, iOS and Android. Drop the converted file in and start reading.

Download Tide Reader

TXT and EPUB; local-first reading, with optional cloud sync and WebDAV.