Round-trip benchmark for technical-document compression and decompression by LLMs. Measures how much
information survives a compress(original) → decompress(compressed) → restored cycle, and how dense
/ readable the intermediate compressed form is.
Size after compression is the compressed text's length as a share of the original. Facts kept is the share of the scenario's fact checklist the judge found in the restored text; invented counts claims the restored text added. Numbers are means over the runs of each cell.
| style | model | runs | size after compression | facts kept | critical facts kept | invented | cost |
|---|---|---|---|---|---|---|---|
| compressed-style | claude-haiku-4-5 | 1 | 0.92 | 100% | 100% | 6 | $0.050 |
| compressed-style | claude-opus-4-7 | 6 | 0.71 | 100% | 100% | 0.8 | $0.062 |
| compressed-style | claude-sonnet-4-6 | 1 | 0.75 | 100% | 100% | 8 | $0.048 |
| concise | claude-opus-4-7 | 6 | 0.92 | 100% | 100% | 5 | $0.050 |
| concise | claude-sonnet-4-6 | 1 | 0.89 | 100% | 100% | 3 | $0.042 |
| concise | claude-haiku-4-5 | 1 | 0.98 | 100% | 100% | 10 | $0.048 |
| glossary-first | claude-sonnet-4-6 | 1 | 0.86 | 100% | 100% | 0 | $0.051 |
| glossary-first | claude-opus-4-7 | 6 | 0.9 | 100% | 100% | 2 | $0.049 |
| glossary-first | claude-haiku-4-5 | 1 | 0.74 | 100% | 100% | 0 | $0.040 |
| naive | claude-opus-4-7 | 4 | 0.63 | 100% | 100% | 7 | $0.046 |
| neutral | claude-opus-4-7 | 2 | 0.63 | 100% | 100% | 2 | $0.045 |
| naive | claude-sonnet-4-6 | 1 | 0.65 | 92.9% | 90% | 11 | $0.044 |
| naive | claude-haiku-4-5 | 1 | 0.72 | 71.4% | 60% | 6 | $0.058 |
required-cli-spawn: The text mentions spawning external CLI processes with isolated HOME but does not name claude, codex, or gemini.import-prefix: Restored text mentions the @bench/ prefix convention but does not state it is defined in deno.json.rejected-node: Text cites npm install and ts-node configuration but does not mention any lockfile policy.rejected-go: Restored text mentions weaker Go fluency and friction with LLM-assisted prompt development, but does not specifically reference 'judge prompts'.required-cli-spawn: Restored text mentions spawning isolated CLI processes but does not name claude/codex/gemini nor mention a per-run isolated HOME directory.| style | model | runs | size after compression | facts kept | critical facts kept | invented | cost |
|---|---|---|---|---|---|---|---|
| compressed-style | claude-opus-4-7 | 1 | 0.71 | 100% | 100% | 0 | $0.157 |
| concise | claude-opus-4-7 | 1 | 0.81 | 100% | 100% | 6 | $0.163 |
| neutral | claude-opus-4-7 | 1 | 0.58 | 85.2% | 86.4% | 4 | $0.129 |
| glossary-first | claude-opus-4-7 | 1 | 0 | 0% | 0% | 0 | $0.000 |
detection-monitor: The 2-minute window is not stated in the restored text; only the >5% threshold is.pagerduty-policy: The escalation policy name 'payments-eu-sev2' is not stated in the restored text.cohorts-unaffected: Text says us-east-1 was not reached, but does not name Cohorts C and D.argo-rollouts: Text states the rollback was via Argo Rollouts and ran cleanly, but does not explicitly state it was 'without manual kubectl intervention'.