Skip to main content
All articles
Format conversion

MOBI/AZW3 → TXT

Converting a novel from MOBI/AZW3 to TXT: what each format is, what gets lost, the pitfalls for long fiction, and the right tools.

Tide Reader Editorial6 min read

What is MOBI/AZW3?

Mobipocket (.mobi) was long the default format Amazon accepted for Kindle delivery. AZW3 (also called KF8) is its successor with fuller HTML5/CSS support, and .azw is an earlier Amazon container. All three reflow and carry navigation and metadata, so they behave as one family. They are proprietary to Amazon's ecosystem, and since 2022 Amazon sends KFX/EPUB instead, so .mobi is mostly a legacy format now.

Extensions.mobi.azw3.azw
Typical useReading on Kindle devices and the Kindle app; one of the most common proprietary formats in existing libraries.

What is TXT?

TXT is plain text: the file holds characters and nothing else, with no fonts, sizes, colours, images or pagination. Line breaks are its only structure. Most Chinese web novels circulate as TXT precisely because it is small, opens anywhere and is easy to copy and update incrementally. The trade-offs are that it carries no chapter structure, cover or metadata, and that mismatched encodings (UTF-8, GBK, UTF-16) produce mojibake — the two most common TXT problems.

Extensions.txt
Typical useChinese web novels, logs, and any text-only content that needs maximum compatibility.

What to watch out for when converting MOBI/AZW3 to TXT

Extracting from the source

.mobi/.azw3 are proprietary Amazon binary containers with no official parsing spec; third-party tools rely on reverse engineering and occasionally fail or drop metadata. AZW3 stores HTML internally, so extracting text is easier than preserving layout.

Writing the target format

TXT stores characters only: every image, font, size, colour, table, footnote and hyperlink is gone, and headings degrade to ordinary lines. You must also choose an encoding (UTF-8 for Chinese, GBK if targeting older devices) — the wrong choice yields immediate mojibake.

Where the two formats interact

  • Illustrations, covers and diagrams in the source are images, and the target stores none — they are all lost, leaving dangling references like "as shown below". If the images matter, target EPUB or PDF instead.
  • Title, author, publisher, language and cover have nowhere to live in the target, so all metadata is lost. If the result goes back into a library, the filename is the only place left to record that information.
  • The source carries real navigation while the target has no concept of chapters, so reading becomes one long scroll. If navigation still matters, keep explicit, uniformly formatted chapter heading lines in the output.
  • Any conversion touching TXT must settle encoding: Chinese TXT files are commonly UTF-8, GBK, GB18030 or BIG5, with or without a BOM. Write UTF-8 without BOM for best compatibility, or GBK for legacy devices. A wrong choice shows up as wholesale mojibake, not as an error.
  • Line endings need normalising too: CRLF on Windows, LF on Unix, and historically CR on classic Mac. Mixing them makes some readers render everything as a single line — if the whole book collapses into one paragraph, check line endings first.

What gets lost

  • Every illustration, cover and diagram
  • Metadata such as title, author and language

What you gain

  • The smallest size and the widest device support

Post-conversion checklist

  1. Convert the first three chapters as a trial and check for lost text, interleaved lines or misplaced blocks before doing the whole book.
  2. Jump to five random spots and search a character name to confirm the text is fully searchable.
  3. Verify encoding and line endings: UTF-8 without BOM, line endings normalised to LF or CRLF.

Converting novels: what is different

Fiction is the most common and the most peculiar kind of e-book: hundreds or thousands of chapters, almost no figures, often nothing but a TXT source, and Chinese sources that frequently carry anti-piracy garbling. These are the points that bite hardest.

  • A long novel can run to hundreds or thousands of chapters, so the table of contents decides whether the result is usable. Normalise chapter headings to one recognisable pattern first — say "Chapter N Title" alone on its own line — or the target will miss chapters, merge them, or collapse the whole book into one.
  • Chinese web novels in TXT show three very different kinds of garbled text. (1) Encoding mismatch — reading UTF-8 as GBK; fixing the encoding restores it. (2) Anti-piracy substitution — the author deliberately replaced characters; it cannot be recovered, so find another source. (3) Full-width/half-width mixing or missing rare glyphs — a font problem. Diagnose which one you have; do not keep re-encoding the third case.
  • Covers, character art and volume-opening illustrations are all dropped, and the library shows a blank cover. Serialised novels often open each volume with an image; afterwards only an empty line is left where it used to be.
  • Volume markers ("Volume One", "Volume Two") usually degrade to ordinary headings, losing the volume/chapter hierarchy. If that matters, promote volume headings one level in an editor afterwards so they are not siblings of chapter titles.

Tools for converting MOBI/AZW3 to TXT

Any of the tools below can handle this direction; pick one by your operating system and how many files you have. Read the previous section first — doing the steps out of order (converting a scanned file before running OCR, for instance) usually means starting over.

Kindle Previewer

FreeDesktop app

Amazon's official tool: it converts EPUB/DOCX to AZW3 and previews the result exactly as a real Kindle would render it. It is the most reliable way to verify how a conversion actually looks on a Kindle, and avoids hand-building the proprietary container.

Platforms: Windows / macOSkdp.amazon.com/en_US/help/topic/G202131100 ↗

Pandoc

FreeCommand line

The Swiss army knife of document conversion, driven by a single command. Its Markdown, DOCX, HTML and EPUB conversions are the highest quality available and it is ideal for authoring-source-to-e-book pipelines. It does not handle MOBI/AZW3 or comic archives.

Platforms: Windows / macOS / Linux (command line)pandoc.org ↗

Calibre

FreeDesktop app

The de-facto standard for e-book conversion, free and open source. It converts EPUB, MOBI, AZW3, PDF, TXT, FB2, DOCX, HTML and nearly everything else, generating tables of contents, fetching metadata and batch-processing whole libraries. The interface is utilitarian, PDF output is mediocre, and the option set takes time to learn.

Platforms: Windows / macOS / Linuxcalibre-ebook.com ↗

Common questions

What is lost when converting MOBI/AZW3 to TXT?

Every illustration, cover and diagram; Metadata such as title, author and language.

Which tool should I use to convert MOBI/AZW3 to TXT?

Recommended: Kindle Previewer, Pandoc, Calibre. Kindle Previewer fits this pair best. Amazon's official tool: it converts EPUB/DOCX to AZW3 and previews the result exactly as a real Kindle would render it. It is the most reliable way to verify how a conversion actually looks on a Kindle, and avoids hand-building the proprietary container.

What should I watch out for when converting MOBI/AZW3 to TXT?

Illustrations, covers and diagrams in the source are images, and the target stores none — they are all lost, leaving dangling references like "as shown below". If the images matter, target EPUB or PDF instead. Title, author, publisher, language and cover have nowhere to live in the target, so all metadata is lost. If the result goes back into a library, the filename is the only place left to record that information. The source carries real navigation while the target has no concept of chapters, so reading becomes one long scroll. If navigation still matters, keep explicit, uniformly formatted chapter heading lines in the output. Any conversion touching TXT must settle encoding: Chinese TXT files are commonly UTF-8, GBK, GB18030 or BIG5, with or without a BOM. Write UTF-8 without BOM for best compatibility, or GBK for legacy devices. A wrong choice shows up as wholesale mojibake, not as an error. Line endings need normalising too: CRLF on Windows, LF on Unix, and historically CR on classic Mac. Mixing them makes some readers render everything as a single line — if the whole book collapses into one paragraph, check line endings first.

Converted it — now what do you read it with?

Tide Reader opens TXT and EPUB, detects chapters on import, lays the text out the way you like it, and syncs your reading position across Windows, macOS, iOS and Android. Drop the converted file in and start reading.

Download Tide Reader

TXT and EPUB; local-first reading, with optional cloud sync and WebDAV.