WETT Benchmark

Which AI models can actually write?

Popular leaderboards obsess over reasoning, maths and code. But writing is the number-one thing people use AI for. So we built WETT — the Writing & Editing Typetone benchmark — to test what matters for content.

Why we built our own benchmark

Newer models keep topping the charts on reasoning and coding — but not on the writing and editing tasks our content engine depends on. Public benchmarks barely measure writing, so we couldn't rely on them to pick the right model. WETT evaluates leading LLMs on the dimensions that decide whether content is actually good.

What WETT measures

Four dimensions of good writing

01

Following writing instructions

Does the model actually do what you asked — including the negative instructions ("don't use these words") that trip most models up?

02

Avoiding LLM-typical language

How well it steers clear of the tell-tale words and phrases that make AI writing sound like AI writing.

03

Stylistic & vocabulary diversity

Whether it varies its expression, or falls back on the same structures and vocabulary again and again.

04

Self-evaluation

How reliably a model can judge the quality of its own output — the basis for automated quality control.

The short version: better at reasoning ≠ better at writing

Across the leading models we tested, many still struggle with negative instructions and stylistic diversity — even as their reasoning scores climb. Which model wins depends on the task. We publish the full methodology, the models tested and the scores in the write-up.

Read the full WETT benchmark

Automated Content Auditing and Compliance at Enterprise Scale.

Enterprise ready integrations, regulatory database and onboarding included.

Join 100.000+ marketing & compliance experts who use Typetone.

European servers Data privacy safe