Benchmark

We saved six reports with 5 editor engines. All 5 rewrote them.

Load a markdown file. Save it without touching anything. Count the lines that changed. On every serialising engine we tested, the answer was never zero.

Method

Six report files were built to look like the markdown people actually receive, each seeded with the constructs that break editors:

FileSimulatesByte-level traps
01 Pentest findingsExternal pentest report: YAML frontmatter, severity tables with alignment, CVSS, HTTP and JSON proofs of concept, task lists, GitHub alerts, details blocks, footnotes, escaped pipesLF
02 AI code auditAgent security review: XML finding tags, emoji severity table, diff blocks, nested quotes, CJK, Greek, Hebrew and Arabic textCRLF + UTF-8 BOM
03 SOC 2 readiness8-column control matrix, exceptions, nested lists with tabs, evidence image with title, trailing spacesMixed CRLF and LF, tabs, trailing whitespace
04 Architecture specDesign RFC: TOML frontmatter, Mermaid, inline and block math, a definition list, reference-style image, HTML reviewer commentLF
05 LLM chat exportChat transcript: nginx configs, tables, Japanese, Chinese, Korean, right-to-left text, literal escape sequencesLF
06 Edge casesSetext headings, every list marker, odd ordered lists, lazy continuation, entities, nested fences, pipe-less tables, three horizontal-rule stylesNo final newline

Each file was loaded and saved with no edits by seven configurations (five serializing engines plus CodeMirror 6 saved two ways), using each engine's documented load and save calls: remark (unified 11 with GFM, frontmatter and math), prosemirror-markdown 1.13, marked to HTML to turndown 7, Milkdown 7.22 (commonmark and GFM presets), Tiptap 3.31 with the official markdown extension, CodeMirror 6 with a naive doc.toString() save, and CodeMirror 6 with the patch-on-save reference technique. These are engines, not the branded apps. Closed GUI apps such as Typora, and Obsidian and VS Code's rendered editors, were not tested directly; Obsidian is built on CodeMirror 6, so it is closer to the CodeMirror rows.

The measure is the percentage of original lines that a line diff marks as removed. 0 means byte-identical. The corpus is pinned by SHA-256 and every package version is pinned in the lockfile.

Result A: save with no edits

Percent of lines changed when each engine loads and saves the file with no edits. Lower is better. 0 is byte-identical.
Reportremarkprosemirror-markdownmarked + turndownMilkdown 7Tiptap 3CodeMirror 6 (naive save)Patch-on-save (reference)
01 Pentest findings (LF)32.5%45.4%44.8%36.2%38.7%00
02 AI code audit (CRLF + BOM)88.7%100%100%88.7%100%100%0
03 SOC 2 readiness (mixed CRLF/LF)64.8%64.8%63.4%64.8%59.2%33.8%0
04 Architecture spec14.3%37.4%38.5%18.7%25.3%00
05 LLM chat export11.2%7.1%16.3%11.2%10.2%00
06 Edge cases (no final newline)36.0%42.1%48.2%38.6%40.4%00
Byte-identical files0/60/60/60/60/64/66/6

Milkdown and Tiptap ran with default presets, without frontmatter, math, footnote or alert extensions. Products built on them often add such plugins, which would fix some construct-level losses but not the reformatting, escaping and line-ending changes that come from re-serialising. The last column is a reference implementation of patch-on-save, the technique AsItIs uses, not the AsItIs app itself.

What they broke

Content lost

  • Tiptap turned the SOC 2 evidence image and the spec's reference-style image into plain text, removed the <details> wrapper, the <dl> definitions and the HTML reviewer comment, and turned the YAML frontmatter into a heading.
  • Milkdown (default presets) dropped the reference-style image entirely and logged an internal error doing it. It turned the opening --- of the YAML frontmatter into a horizontal rule, so the metadata became body text, and stripped the BOM.
  • marked + turndown removed all four <finding> tags with their severity attributes from the AI audit, so the severity metadata is gone, along with <details>, <dl>, kbd, mark, sup and the reviewer comment. Task items became plain bullets.
  • prosemirror-markdown has no table node in its default schema: every table in the pentest report and the SOC 2 control matrix was flattened into one paragraph of pipes. It also escaped footnotes and merged their definitions, escaped task checkboxes, and broke both frontmatter blocks and the math block.
  • remark kept most constructs, with frontmatter, GFM and math plugins enabled, but still changed 11% to 89% of lines.

Meaning silently changed

  • Every serialising engine escaped the GitHub alert: > [!CAUTION] became > \[!CAUTION], so GitHub renders an ordinary quote instead of the red Caution box. In a security report, that box is the handling notice.
  • Tiptap escaped footnotes ([^idor] became \[^idor\]), so they render as literal text, and double-escaped entities: &nbsp; became &amp;nbsp;, which renders as the literal text &nbsp;.
  • Line endings: every engine except patch-on-save changed the line endings of the CRLF file and flattened the mixed file. remark and Milkdown also stripped the BOM. In a Windows repository that shows up as every line changed in git diff.

Noise

  • Tables re-padded (remark, Milkdown, Tiptap). List markers normalised: + and - became * in remark, prosemirror-markdown and Milkdown, and -    in turndown. The 1. 1. 1. list was renumbered 1. 2. 3. by all five.
  • Hard breaks converted: two trailing spaces became a backslash in remark, prosemirror-markdown and Milkdown, and the reverse in turndown and Tiptap.

The naive CodeMirror 6 save is byte-perfect on LF files, but rewrites 100% of the CRLF file and a third of the mixed file, because the editor normalises line breaks to LF. A source editor needs a save layer too. An editor that sets a lineSeparator would keep a uniformly CRLF file; a mixed file still needs per-line handling.

Result B: real edits

Edits issued the way rendered-view widgets would issue them, as minimal source patches, through the patch-on-save reference.

FileEditsResultA naive save would change
01 Pentest findings (LF)Tick a task, edit a table cell, insert a table row, fix a word3 changed, 1 added3 lines
02 AI code audit (CRLF + BOM)Edit a table cell, insert a 2-line paragraph1 changed, 3 added. BOM kept, new lines written as CRLF (115 to 118 CRLF)115 lines
03 SOC 2 (mixed CRLF/LF)Tick a task, edit a cell in an 8-column table2 changed. The 24 CRLF and 47 LF lines stay as they were25 lines
06 Edge cases (no final newline)Edit a cell in a pipe-less table, edit the last line2 changed. Still no final newline2 lines

Every byte outside the edited lines (BOM, per-line endings, trailing spaces, tabs, final newline state) is copied from the original.

Result C: speed

Operation (ms, one run each)1 MB5 MB20 MB
lezer full parse (CodeMirror's markdown parser)58217873
Patch-on-save: open, 3 edits, save952222
markdown-it render653811,216
marked render51287906
prosemirror-markdown parse39184skipped
remark parse1,05818,747skipped
Milkdown load and save20,596skippedskipped
Tiptap load and save111,611skippedskipped

Node 26 on Apple silicon (M5). Milkdown and Tiptap ran in jsdom, which is slower than a browser, so treat their absolute times as pessimistic; the gap is the signal. Timings are single runs and vary by about a factor of two between runs. The round-trip, construct and edit results are deterministic.

Limits of this benchmark

  • Six files, deliberately dense with hard constructs. Plain prose with a few headings round-trips much better in every engine (see file 05, where engines changed 7 to 16% of lines).
  • "Lines changed" counts formatting changes too. Re-padding a table is not content loss, but it makes review and git diff noisy.
  • UTF-8 only. UTF-16 and legacy encodings are not tested.

Reproduce it

git clone https://github.com/Katta041/markdown-roundtrip-benchmark.git
cd markdown-roundtrip-benchmark
npm ci
npm run bench -- --quick   # a few seconds: round trip, edits, 1 MB timings
npm run bench              # full run, about 3 minutes: adds 5 MB and 20 MB timings

Requires Node 24 or later. The run writes results/RESULTS.md and the exact bytes each engine saved to results/output/<engine>/, so you can diff it yourself. To add an engine or correct an unfair configuration, open a pull request: an adapter is one object with a roundTrip function. Everything is on GitHub, MIT licensed.

Results as published in the repository on 2026-10-01.