TXT → EPUB
Converting a novel from TXT to EPUB: what each format is, what gets lost, the pitfalls for long fiction, and the right tools.
What is TXT?
TXT is plain text: the file holds characters and nothing else, with no fonts, sizes, colours, images or pagination. Line breaks are its only structure. Most Chinese web novels circulate as TXT precisely because it is small, opens anywhere and is easy to copy and update incrementally. The trade-offs are that it carries no chapter structure, cover or metadata, and that mismatched encodings (UTF-8, GBK, UTF-16) produce mojibake — the two most common TXT problems.
.txtWhat is EPUB?
EPUB is the open e-book standard maintained by the International Digital Publishing Forum. Technically it is a ZIP archive holding XHTML content, CSS, images and an OPF manifest that defines reading order. Its defining trait is that it reflows: content and presentation are separate, so a reader re-lays-out the text for the current screen and font size. That is why one file looks right on a phone, a tablet and a desktop, and why EPUB is the most portable format outside the Kindle ecosystem.
.epubWhat to watch out for when converting TXT to EPUB
Extracting from the source
TXT carries no chapter markers — a heading is just a line that looks like one. Before converting, confirm the text has been split correctly (a regex matching chapter patterns), otherwise the whole book becomes one un-navigable block inside EPUB or Kindle.
Writing the target format
EPUB is not just a ZIP: it must contain mimetype (as the first, uncompressed entry), META-INF/container.xml, the content OPF and a nav document. Miss one and readers report a corrupt file, so prefer a mature converter over hand-building the archive.
Where the two formats interact
- The source contains no images, so the result has no cover or illustrations. EPUB/Kindle library views then show a blank placeholder — add a cover manually afterwards or the title is unrecognisable on a shelf.
- The source carries no metadata at all (TXT is the classic case) while the target expects a title and author. Converters fall back to the filename, so libraries fill up with "Untitled" or titles still carrying the .txt suffix — fill them in during conversion.
- Chapter headings in the source are ordinary lines while the target demands a real table of contents. Converters try to match patterns such as "Chapter N"; when the match is poor the TOC is garbled or the book has a single chapter. Check that TOC entries match the actual chapter count.
- Going from a single text file to a structured archive adds information: the tool must generate the manifest, the navigation document and metadata for you. The quality of those generated parts — accurate chapter splitting, correct language tags — decides whether the result is usable.
- Any conversion touching TXT must settle encoding: Chinese TXT files are commonly UTF-8, GBK, GB18030 or BIG5, with or without a BOM. Write UTF-8 without BOM for best compatibility, or GBK for legacy devices. A wrong choice shows up as wholesale mojibake, not as an error.
- Line endings need normalising too: CRLF on Windows, LF on Unix, and historically CR on classic Mac. Mixing them makes some readers render everything as a single line — if the whole book collapses into one paragraph, check line endings first.
What gets lost
- Both sides are structurally similar; only fine typographic detail is lost
What you gain
- A real table of contents with chapter jumps
- A place for title, author and cover metadata
- Full support in mainstream readers and libraries
Post-conversion checklist
- Convert the first three chapters as a trial and check for lost text, interleaved lines or misplaced blocks before doing the whole book.
- Check chapter splitting: the TOC entry count should match the real chapter count.
- Jump to five random spots and search a character name to confirm the text is fully searchable.
- Fill in title, author and language — a wrong language tag breaks typesetting such as Chinese line breaking.
- Open the result in two different readers to confirm the TOC and cover — EPUB is the strictest about structure.
Converting novels: what is different
Fiction is the most common and the most peculiar kind of e-book: hundreds or thousands of chapters, almost no figures, often nothing but a TXT source, and Chinese sources that frequently carry anti-piracy garbling. These are the points that bite hardest.
- A long novel can run to hundreds or thousands of chapters, so the table of contents decides whether the result is usable. Normalise chapter headings to one recognisable pattern first — say "Chapter N Title" alone on its own line — or the target will miss chapters, merge them, or collapse the whole book into one.
- Chinese web novels in TXT show three very different kinds of garbled text. (1) Encoding mismatch — reading UTF-8 as GBK; fixing the encoding restores it. (2) Anti-piracy substitution — the author deliberately replaced characters; it cannot be recovered, so find another source. (3) Full-width/half-width mixing or missing rare glyphs — a font problem. Diagnose which one you have; do not keep re-encoding the third case.
- Volume markers ("Volume One", "Volume Two") usually degrade to ordinary headings, losing the volume/chapter hierarchy. If that matters, promote volume headings one level in an editor afterwards so they are not siblings of chapter titles.
Tools for converting TXT to EPUB
Any of the tools below can handle this direction; pick one by your operating system and how many files you have. Read the previous section first — doing the steps out of order (converting a scanned file before running OCR, for instance) usually means starting over.
Sigil
FreeDesktop appA visual EPUB editor that lets you edit the internal XHTML/CSS directly, fix the TOC and repair metadata. It suits polishing an EPUB after conversion — wrong chapter split, missing cover, layout to adjust. It does not convert between formats, only edits EPUB.
Calibre
FreeDesktop appThe de-facto standard for e-book conversion, free and open source. It converts EPUB, MOBI, AZW3, PDF, TXT, FB2, DOCX, HTML and nearly everything else, generating tables of contents, fetching metadata and batch-processing whole libraries. The interface is utilitarian, PDF output is mediocre, and the option set takes time to learn.
SumatraPDF
FreeReaderA lightweight reader for PDF, EPUB, MOBI, CBZ and more that can save PDF content as text. It is a fast way to check whether text extracts cleanly — if the selection comes out garbled, OCR is needed before converting.
Common questions
What is lost when converting TXT to EPUB?
Both sides are structurally similar; only fine typographic detail is lost.
Which tool should I use to convert TXT to EPUB?
Recommended: Sigil, Calibre, SumatraPDF. Sigil fits this pair best. A visual EPUB editor that lets you edit the internal XHTML/CSS directly, fix the TOC and repair metadata. It suits polishing an EPUB after conversion — wrong chapter split, missing cover, layout to adjust. It does not convert between formats, only edits EPUB.
What should I watch out for when converting TXT to EPUB?
The source contains no images, so the result has no cover or illustrations. EPUB/Kindle library views then show a blank placeholder — add a cover manually afterwards or the title is unrecognisable on a shelf. The source carries no metadata at all (TXT is the classic case) while the target expects a title and author. Converters fall back to the filename, so libraries fill up with "Untitled" or titles still carrying the .txt suffix — fill them in during conversion. Chapter headings in the source are ordinary lines while the target demands a real table of contents. Converters try to match patterns such as "Chapter N"; when the match is poor the TOC is garbled or the book has a single chapter. Check that TOC entries match the actual chapter count. Going from a single text file to a structured archive adds information: the tool must generate the manifest, the navigation document and metadata for you. The quality of those generated parts — accurate chapter splitting, correct language tags — decides whether the result is usable. Any conversion touching TXT must settle encoding: Chinese TXT files are commonly UTF-8, GBK, GB18030 or BIG5, with or without a BOM. Write UTF-8 without BOM for best compatibility, or GBK for legacy devices. A wrong choice shows up as wholesale mojibake, not as an error. Line endings need normalising too: CRLF on Windows, LF on Unix, and historically CR on classic Mac. Mixing them makes some readers render everything as a single line — if the whole book collapses into one paragraph, check line endings first.
Converted it — now what do you read it with?
Tide Reader opens TXT and EPUB, detects chapters on import, lays the text out the way you like it, and syncs your reading position across Windows, macOS, iOS and Android. Drop the converted file in and start reading.
Download Tide ReaderTXT and EPUB; local-first reading, with optional cloud sync and WebDAV.
