Which AI models can actually write?
Popular leaderboards obsess over reasoning, maths and code. But writing is the number-one thing people use AI for. So we built WETT — the Writing & Editing Typetone benchmark — to test what matters for content.
Why we built our own benchmark
Newer models keep topping the charts on reasoning and coding — but not on the writing and editing tasks our content engine depends on. Public benchmarks barely measure writing, so we couldn't rely on them to pick the right model. WETT evaluates leading LLMs on the dimensions that decide whether content is actually good.
Four dimensions of good writing
Following writing instructions
Does the model actually do what you asked — including the negative instructions ("don't use these words") that trip most models up?
Avoiding LLM-typical language
How well it steers clear of the tell-tale words and phrases that make AI writing sound like AI writing.
Stylistic & vocabulary diversity
Whether it varies its expression, or falls back on the same structures and vocabulary again and again.
Self-evaluation
How reliably a model can judge the quality of its own output — the basis for automated quality control.
The short version: better at reasoning ≠ better at writing
Across the leading models we tested, many still struggle with negative instructions and stylistic diversity — even as their reasoning scores climb. Which model wins depends on the task. We publish the full methodology, the models tested and the scores in the write-up.
Read the full WETT benchmarkAutomated Content Auditing and Compliance at Enterprise Scale.
Enterprise ready integrations, regulatory database and onboarding included.
Join 100.000+ marketing & compliance experts who use Typetone.