Small benchmarks of language models and coding agents: what they cost in tokens, how they write, what they keep and what they draw. Each page shows the latest results and how they were measured.
How many tokens each model's tokenizer spends on the same text, the Universal Declaration of Human Rights, in dozens of languages.
41 models, 52 texts.
Whether a coding agent's messages stay understandable to the user: its own terms, the user's context, one language.
Best on all 4 scenarios: claude:claude-opus-5-5:high, score 80%, solved 2/4.
How much a one-line "be brief" modifier shortens model answers to 10 concrete questions, and whether the answers stay right.
"Be brief" cuts words 4–10× and every reply stays correct.
How much of a technical document survives when an LLM compresses it and another call restores it.
Best: claude-haiku-4-5 with the compressed-style prompt keeps 100% of facts.
Twelve text-to-image prompts that expose model limits: text on images, counting, layout, consistency.
6 image models on 12 hard prompts.
One Russian prompt for a children's alphabet poster with a DevOps word for every letter, and what three models drew.
3 posters to compare side by side.