HTML → Markdown
Converting a novel from HyperText (.html, .htm, .xhtml, .xml, .mhtml) to Markdown (.md) is a routine step in practice — Markdown is plain text, so any editor opens it and it versions cleanly. Going from HyperText to Markdown does raise extra questions worth thinking about: you lose metadata such as title, author and language. Below we explain in detail how HyperText becomes Markdown and what to watch out for, and the page ends with a converter that does the job right in your browser — nothing is uploaded, the file never leaves your machine.
What is HTML?
HTML and XHTML are the structural languages of the web, and XHTML is also what EPUB uses internally to hold its content — so converting between HTML and EPUB is essentially a matter of swapping the shell and how assets are packaged. .mhtml goes further and packs a page together with its images and styles into one file, which suits preserving online pages that may disappear. HTML's strength is that it is universal and opens in any browser; its weakness is that styles and images are usually external links, so shipping one self-contained file means inlining them first. Read the HTML format guide
| Extensions | .html.htm.xhtml.xml.mhtml |
|---|---|
| Typical use | Captured web content, EPUB internal content, and a general intermediate representation for other formats. |
What is Markdown?
Markdown is plain text plus markers (# headings, ** bold, | tables) — effectively a superset of .txt. Any text editor opens it, and it versions cleanly in Git. Its expressiveness is deliberately limited: no pagination, no headers or footers, no precise control over type size or position, and no complex layouts. Those trade-offs make Markdown an excellent authoring and conversion source, but as a final reading format its experience depends entirely on which app opens it. Read the Markdown format guide
| Extensions | .md |
|---|---|
| Typical use | Authoring, notes and technical docs; a clean intermediate format before converting to EPUB or TXT. |
How to convert HTML to Markdown
All structure lives in the tags, but real pages mix in navigation, ads and footers that a naive conversion carries along. Extract the content container first (a readability pass), and note that <img> usually points outside the file — images must be downloaded and embedded before producing a single-file format.
Markdown output keeps text and only the most basic hierarchy (# headings, paragraphs, simple lists). Pagination, headers and footers, comments, tracked changes, precise type sizes and text wrapping are all gone — it is a content archive, not a layout archive.
Metadata in the HTML source — title, author, publisher, language, table of contents, cover — has nowhere to live in the target, so converting to Markdown loses all of it.
Convert HTML to Markdown
HTML → MARKDOWN
The conversion runs entirely in your browser — your file is never uploaded.
A novel reader that syncs your progress across every platform
Tide Reader opens TXT and EPUB, detects chapters on import, lays the text out the way you like it, and syncs your reading position across Windows, macOS, iOS and Android. Drop the converted file in and start reading.
Download Tide ReaderTXT and EPUB; local-first reading, with optional cloud sync and WebDAV.
