Skip to main content
All articles
Format conversion

TXT → PDF

Converting a novel from TXT to PDF: what each format is, what gets lost, the pitfalls for long fiction, and the right tools.

Tide Reader Editorial6 min read

What is TXT?

TXT is plain text: the file holds characters and nothing else, with no fonts, sizes, colours, images or pagination. Line breaks are its only structure. Most Chinese web novels circulate as TXT precisely because it is small, opens anywhere and is easy to copy and update incrementally. The trade-offs are that it carries no chapter structure, cover or metadata, and that mismatched encodings (UTF-8, GBK, UTF-16) produce mojibake — the two most common TXT problems.

Extensions.txt
Typical useChinese web novels, logs, and any text-only content that needs maximum compatibility.

What is PDF?

PDF is Adobe's fixed-layout document format. It records what to draw at which coordinate on a page rather than which paragraph a sentence belongs to. The layout is therefore identical everywhere, which suits contracts, papers and scans. The cost is that PDF is inherently hostile to reflow: line breaks are usually visual, extracted text often carries hard returns and scrambled column order, and scanned PDFs contain images only, with no text layer at all.

Extensions.pdf
Typical useDocuments, papers, contracts and scans where the layout must stay exact; also a print and archival format.

What to watch out for when converting TXT to PDF

Extracting from the source

TXT carries no chapter markers — a heading is just a line that looks like one. Before converting, confirm the text has been split correctly (a regex matching chapter patterns), otherwise the whole book becomes one un-navigable block inside EPUB or Kindle.

Writing the target format

Producing PDF freezes the content: once generated it cannot reflow, so readers cannot change font size or line width and long text is painful on small screens. For extended reading PDF is usually the wrong target — it suits printing, archiving and layout-critical delivery.

Where the two formats interact

  • Reflow is lost: the target freezes the layout, so readers can no longer change font size or line width. Long novels become zoom-and-pan on small screens — the most common reason a conversion ends up harder to read. Unless you need print or layout-critical delivery, keep a reflowable copy for reading.
  • The source carries no metadata at all (TXT is the classic case) while the target expects a title and author. Converters fall back to the filename, so libraries fill up with "Untitled" or titles still carrying the .txt suffix — fill them in during conversion.
  • Any conversion touching TXT must settle encoding: Chinese TXT files are commonly UTF-8, GBK, GB18030 or BIG5, with or without a BOM. Write UTF-8 without BOM for best compatibility, or GBK for legacy devices. A wrong choice shows up as wholesale mojibake, not as an error.
  • Line endings need normalising too: CRLF on Windows, LF on Unix, and historically CR on classic Mac. Mixing them makes some readers render everything as a single line — if the whole book collapses into one paragraph, check line endings first.

What gets lost

  • Adjustable font size and screen-based reflow
  • Text reflow (a generated PDF is fixed for good)

What you gain

  • A place for title, author and cover metadata
  • Identical layout everywhere, ideal for printing

Post-conversion checklist

  1. Convert the first three chapters as a trial and check for lost text, interleaved lines or misplaced blocks before doing the whole book.
  2. Check chapter splitting: the TOC entry count should match the real chapter count.
  3. Jump to five random spots and search a character name to confirm the text is fully searchable.
  4. Fill in title, author and language — a wrong language tag breaks typesetting such as Chinese line breaking.
  5. Check margins and that the Chinese font is embedded, otherwise it renders as boxes on another machine.

Converting novels: what is different

Fiction is the most common and the most peculiar kind of e-book: hundreds or thousands of chapters, almost no figures, often nothing but a TXT source, and Chinese sources that frequently carry anti-piracy garbling. These are the points that bite hardest.

  • The target has no notion of a table of contents, and not being able to jump to a chapter is the worst part of reading something long. If you will keep reading it, keep a navigable copy as well (EPUB or Kindle) and treat this one as the distribution or archive copy.
  • Chinese web novels in TXT show three very different kinds of garbled text. (1) Encoding mismatch — reading UTF-8 as GBK; fixing the encoding restores it. (2) Anti-piracy substitution — the author deliberately replaced characters; it cannot be recovered, so find another source. (3) Full-width/half-width mixing or missing rare glyphs — a font problem. Diagnose which one you have; do not keep re-encoding the third case.
  • Once a long novel is in a fixed layout, reading on a phone means zooming and panning. Fiction is the worst fit for fixed layout of any genre: almost no figures, all continuous text. Unless you are printing or submitting it, prefer a reflowable format for novels.
  • Volume markers ("Volume One", "Volume Two") usually degrade to ordinary headings, losing the volume/chapter hierarchy. If that matters, promote volume headings one level in an editor afterwards so they are not siblings of chapter titles.

Tools for converting TXT to PDF

Any of the tools below can handle this direction; pick one by your operating system and how many files you have. Read the previous section first — doing the steps out of order (converting a scanned file before running OCR, for instance) usually means starting over.

ocrmypdf / Tesseract

FreeCommand line

An open-source way to add a text layer (OCR) to scanned PDFs, with Chinese support. Scanned files must go through this before any text-extraction conversion, otherwise the output is blank. Expect to proofread the recognised characters.

Platforms: Windows / macOS / Linux (command line)ocrmypdf.readthedocs.io ↗

Stirling PDF

FreeOnline service

An open-source PDF toolbox that splits, merges, rotates, compresses and extracts text in the browser, and can be self-hosted so files never leave your network. Good for the PDF side; for reflowable output Calibre remains the better choice.

Platforms: Browser (self-hostable)stirlingpdf.io ↗

Calibre

FreeDesktop app

The de-facto standard for e-book conversion, free and open source. It converts EPUB, MOBI, AZW3, PDF, TXT, FB2, DOCX, HTML and nearly everything else, generating tables of contents, fetching metadata and batch-processing whole libraries. The interface is utilitarian, PDF output is mediocre, and the option set takes time to learn.

Platforms: Windows / macOS / Linuxcalibre-ebook.com ↗

Common questions

What is lost when converting TXT to PDF?

Adjustable font size and screen-based reflow; Text reflow (a generated PDF is fixed for good).

Which tool should I use to convert TXT to PDF?

Recommended: ocrmypdf / Tesseract, Stirling PDF, Calibre. ocrmypdf / Tesseract fits this pair best. An open-source way to add a text layer (OCR) to scanned PDFs, with Chinese support. Scanned files must go through this before any text-extraction conversion, otherwise the output is blank. Expect to proofread the recognised characters.

What should I watch out for when converting TXT to PDF?

Reflow is lost: the target freezes the layout, so readers can no longer change font size or line width. Long novels become zoom-and-pan on small screens — the most common reason a conversion ends up harder to read. Unless you need print or layout-critical delivery, keep a reflowable copy for reading. The source carries no metadata at all (TXT is the classic case) while the target expects a title and author. Converters fall back to the filename, so libraries fill up with "Untitled" or titles still carrying the .txt suffix — fill them in during conversion. Any conversion touching TXT must settle encoding: Chinese TXT files are commonly UTF-8, GBK, GB18030 or BIG5, with or without a BOM. Write UTF-8 without BOM for best compatibility, or GBK for legacy devices. A wrong choice shows up as wholesale mojibake, not as an error. Line endings need normalising too: CRLF on Windows, LF on Unix, and historically CR on classic Mac. Mixing them makes some readers render everything as a single line — if the whole book collapses into one paragraph, check line endings first.

Converted it — now what do you read it with?

Tide Reader opens TXT and EPUB, detects chapters on import, lays the text out the way you like it, and syncs your reading position across Windows, macOS, iOS and Android. Drop the converted file in and start reading.

Download Tide Reader

TXT and EPUB; local-first reading, with optional cloud sync and WebDAV.