LLM benchmarks

Small benchmarks of language models and coding agents: what they cost in tokens, how they write, what they keep and what they draw. Each page shows the latest results and how they were measured.

Tokenizers: tokens per language

How many tokens each model's tokenizer spends on the same text, the Universal Declaration of Human Rights, in dozens of languages.

41 models, 52 texts.

Drift: does an agent stay understandable

Whether a coding agent's messages stay understandable to the user: its own terms, the user's context, one language.

Best on all 4 scenarios: claude:claude-opus-5-5:high, score 80%, solved 2/4.

Brevity: what "be brief" does to answers

How much a one-line "be brief" modifier shortens model answers to 10 concrete questions, and whether the answers stay right.

"Be brief" cuts words 4–10× and every reply stays correct.

Compression: what survives compress → decompress

How much of a technical document survives when an LLM compresses it and another call restores it.

Best: claude-haiku-4-5 with the compressed-style prompt keeps 100% of facts.

Images-Hard: hard prompts for image models

Twelve text-to-image prompts that expose model limits: text on images, counting, layout, consistency.

6 image models on 12 hard prompts.

DevOps alphabet: one prompt, three image models

One Russian prompt for a children's alphabet poster with a DevOps word for every letter, and what three models drew.

3 posters to compare side by side.