Skip to main content
All articles
Format conversion

HTML → CBZ/CBR

Converting a novel from HTML to CBZ/CBR: what each format is, what gets lost, the pitfalls for long fiction, and the right tools.

Tide Reader Editorial6 min read

What is HTML?

HTML / XHTML is the structural language of the web and is literally the content format inside an EPUB; .mhtml (MHTML) packs a page together with all its resources (images, CSS) into one file, which is what "save as a single page" produces. XML is the general markup language that formats such as FB2 and OPF are themselves written in. The whole family is tagged text where the tags carry the structure, so how much structure survives a conversion depends on how well the tool understands them.

Extensions.html.htm.xhtml.xml.mhtml
Typical useCaptured web content, EPUB internal content, and a general intermediate representation for other formats.

What is CBZ/CBR?

A comic book archive is not a real format but a folder of images packed in order: .cbz is ZIP, .cbr is RAR, .cbt is TAR and .cb7 is 7z, renamed so readers recognise it. There is no text layer and no typography — each page is a single full-page image, and reading order comes from filename sorting (which is why unpadded names like 1.jpg, 2.jpg … 10.jpg put page 10 right after page 1).

Extensions.cbr.cbz.cbt.cb7
Typical useComics, art books and books scanned to images; fundamentally an image bundle rather than a text book.

What to watch out for when converting HTML to CBZ/CBR

Extracting from the source

All structure lives in the tags, but real pages mix in navigation, ads and footers that a naive conversion carries along. Extract the content container first (a readability pass), and note that <img> usually points outside the file — images must be downloaded and embedded before producing a single-file format.

Writing the target format

Converting to a comic archive really just packs images under ordered names; all text is discarded. Zero-pad the filenames (001.jpg, 002.jpg) so sorting works, and note that .cbr needs RAR tools while most readers handle .cbz (ZIP) best.

Where the two formats interact

  • Reflow is lost: the target freezes the layout, so readers can no longer change font size or line width. Long novels become zoom-and-pan on small screens — the most common reason a conversion ends up harder to read. Unless you need print or layout-critical delivery, keep a reflowable copy for reading.
  • The target is image-only: all text in the source is discarded (or survives only as pixels). If your goal is content that can be searched, copied or read aloud, this direction is wrong — pick a text-preserving target instead.
  • Going from a single text file to a structured archive adds information: the tool must generate the manifest, the navigation document and metadata for you. The quality of those generated parts — accurate chapter splitting, correct language tags — decides whether the result is usable.

What gets lost

  • Adjustable font size and screen-based reflow
  • The searchable, copyable text layer
  • CSS presentation: fonts, sizes, colours, indents and spacing

What you gain

  • A page-ordered image bundle that comic readers page through

Post-conversion checklist

  1. Convert the first three chapters as a trial and check for lost text, interleaved lines or misplaced blocks before doing the whole book.
  2. Check chapter splitting: the TOC entry count should match the real chapter count.
  3. Confirm illustration count and placement — check that images are not all dumped at the end of a chapter.
  4. Fill in title, author and language — a wrong language tag breaks typesetting such as Chinese line breaking.

Converting novels: what is different

Fiction is the most common and the most peculiar kind of e-book: hundreds or thousands of chapters, almost no figures, often nothing but a TXT source, and Chinese sources that frequently carry anti-piracy garbling. These are the points that bite hardest.

  • The target has no notion of a table of contents, and not being able to jump to a chapter is the worst part of reading something long. If you will keep reading it, keep a navigable copy as well (EPUB or Kindle) and treat this one as the distribution or archive copy.
  • Once a long novel is in a fixed layout, reading on a phone means zooming and panning. Fiction is the worst fit for fixed layout of any genre: almost no figures, all continuous text. Unless you are printing or submitting it, prefer a reflowable format for novels.
  • Volume markers ("Volume One", "Volume Two") usually degrade to ordinary headings, losing the volume/chapter hierarchy. If that matters, promote volume headings one level in an editor afterwards so they are not siblings of chapter titles.

Tools for converting HTML to CBZ/CBR

Any of the tools below can handle this direction; pick one by your operating system and how many files you have. Read the previous section first — doing the steps out of order (converting a scanned file before running OCR, for instance) usually means starting over.

Kindle Comic Converter

FreeDesktop app

A converter built for comic archives: CBZ/CBR/CB7 in, output tuned for Kindle, Kobo and similar devices. It handles image scaling, greyscale, double-page spreads and zero-padding of filenames, all of which are tedious by hand.

Platforms: Windows / macOS / Linuxgithub.com/ciromattia/kcc ↗

Calibre

FreeDesktop app

The de-facto standard for e-book conversion, free and open source. It converts EPUB, MOBI, AZW3, PDF, TXT, FB2, DOCX, HTML and nearly everything else, generating tables of contents, fetching metadata and batch-processing whole libraries. The interface is utilitarian, PDF output is mediocre, and the option set takes time to learn.

Platforms: Windows / macOS / Linuxcalibre-ebook.com ↗

SumatraPDF

FreeReader

A lightweight reader for PDF, EPUB, MOBI, CBZ and more that can save PDF content as text. It is a fast way to check whether text extracts cleanly — if the selection comes out garbled, OCR is needed before converting.

Platforms: Windowswww.sumatrapdfreader.org ↗

Common questions

What is lost when converting HTML to CBZ/CBR?

Adjustable font size and screen-based reflow; The searchable, copyable text layer; CSS presentation: fonts, sizes, colours, indents and spacing.

Which tool should I use to convert HTML to CBZ/CBR?

Recommended: Kindle Comic Converter, Calibre, SumatraPDF. Kindle Comic Converter fits this pair best. A converter built for comic archives: CBZ/CBR/CB7 in, output tuned for Kindle, Kobo and similar devices. It handles image scaling, greyscale, double-page spreads and zero-padding of filenames, all of which are tedious by hand.

What should I watch out for when converting HTML to CBZ/CBR?

Reflow is lost: the target freezes the layout, so readers can no longer change font size or line width. Long novels become zoom-and-pan on small screens — the most common reason a conversion ends up harder to read. Unless you need print or layout-critical delivery, keep a reflowable copy for reading. The target is image-only: all text in the source is discarded (or survives only as pixels). If your goal is content that can be searched, copied or read aloud, this direction is wrong — pick a text-preserving target instead. Going from a single text file to a structured archive adds information: the tool must generate the manifest, the navigation document and metadata for you. The quality of those generated parts — accurate chapter splitting, correct language tags — decides whether the result is usable.

Converted it — now what do you read it with?

Tide Reader opens TXT and EPUB, detects chapters on import, lays the text out the way you like it, and syncs your reading position across Windows, macOS, iOS and Android. Drop the converted file in and start reading.

Download Tide Reader

TXT and EPUB; local-first reading, with optional cloud sync and WebDAV.