Five levels, from slop to work that sells
Score any piece from 0 to 4. Each level has one test, and one piece of copy below shows what changes at each step.
The scale and the test for each level
| Level | Name | The test |
|---|---|---|
| 0 | Slop | Was it sent as the model wrote it, with no check? |
| 1 | Clean | Are the obvious machine signs gone: the stock words, the dashes, the three-part lists? |
| 2 | Correct | Has a named person checked each fact and opened each source? |
| 3 | Clear and specific | Is it written to one reader, in plain words, with facts where the adjectives were? |
| 4 | Distinct and effective | Does it hold an angle and a fact only you have, and was it tested on readers? |
A piece must pass each lower level first. A strong idea with a wrong number fails level 2.
To move up: from 1 to 2, open each source. From 2 to 3, cut each sentence over 25 words and replace each adjective with a number or a name. From 3 to 4, add the fact only you hold and test two versions.
One piece of copy at each level
This example is invented. The agency, the client and the figures are not real.
Level 0. The model's first output, sent as it came.
In today's fast-paced digital landscape, brands need more than just great content — they need a strategic partner. At Northfield, we don't just create campaigns, we craft experiences that resonate, engage and convert. Ready to unlock your brand's potential? Let's talk.
Level 1. The stock words and the dash are gone. It still says what any agency could say.
Brands need content that works and a partner who understands strategy. At Northfield we plan and make campaigns that reach the right people and bring results. Contact us to discuss your brand.
Level 2. Each claim is now a fact that someone checked.
Northfield plans and makes campaigns for retail brands. In 2026 we ran 14 campaigns for nine clients. Contact us to discuss your brand.
Level 3. One reader, their problem, one number, one request.
You run marketing for a garden-centre chain, and spring is 12 weeks away. Last spring we rewrote 40 product emails for one chain, and its email sales rose 18% in six weeks. We can show you the before and after in 20 minutes.
Level 4. An angle the reader has not heard, a fact only this agency holds, and a test.
Your customers buy compost in March when it rains on Saturday. Your email has little to do with it. Last spring we tied one garden chain's emails to the weekend forecast, and email sales rose 18% in six weeks. We tested the weather subject line against the old one first, and it got 31% more opens. Give us 20 minutes and we will show you how.
Level 0 is common and it costs the reader time
- 40% of 1,150 US desk workers said they had received AI "workslop" in the past month. Each case cost them about two hours (BetterUp Labs and Stanford Social Media Lab, September 2025; a vendor co-wrote the survey).
- In the working paper of one experiment with 444 professionals, 68% of those given ChatGPT submitted its first output with no edit (Noy and Zhang, 2023).
Level 1 is where most "humanised" copy stops
Some of the famous words are fading. In arXiv abstracts, "delve" began to fall from April 2024, soon after it was named as a ChatGPT word (Geng and Trotta, 2025). The sentence patterns are harder to remove. The grammar and rhetoric of model text differ from human text, and the difference persists in larger models (Reinhart and others, PNAS 2025).
Level 2 needs a named person
- Associated Press: "Any output from a generative AI tool should be treated as unvetted source material." (AP)
- BBC: AI use "must include active human editorial oversight and approval" (BBC editorial guidance).
- The US agency body's policy template: "Our team members will always review, modify and edit any text or images generated by AI before we include them in a work product" (4A's).
The UK Advertising Association guide (2026) sets the depth of review by risk:
| Risk tier | Examples | Review the guide requires |
|---|---|---|
| High | Regulated sectors, vulnerable groups, claims about health, safety or product efficacy | "senior executive approval and multiple review stages" |
| Limited | Authorised and disclosed use of a real person's likeness | "managerial approval and expert review" |
| Minimal | General AI-generated creative with no likeness | "automated screening and periodic audits" |
Source: Advertising Association best practice guide.
Level 3 gets more response
- Simpler headlines got more clicks in 7,371 Washington Post tests and 22,664 Upworthy tests. General readers picked the simple headline at an odds ratio of 2.83, and journalists showed no preference (Shulman, Markowitz and Rogers, Science Advances, 2024).
- A 49-word email got a 4.8% response. A 127-word version of the same request got 2.7% (Rogers and Lasky-Fink, 2023, 7,002 US school board members).
- The UK government standard gives two limits: split a sentence of more than 25 words, and keep a paragraph to five sentences (GOV.UK).
Level 4 is where the money is
- Creative quality drove 49% of the incremental sales from advertising in nearly 450 campaigns for packaged consumer goods (NCSolutions, vendor study, 2023).
- The most common response to the average TV ad is no feeling: 52% of responses in the UK and 47% in the US (System1, The Extraordinary Cost of Dull, vendor study, 2024).
- Level 4 needs a test because a model alone cannot pick the winner. In 17,681 headline tests, the best model-only method was right about 47% of the time, which the authors call marginally better than random (Ye, Yoganarasimhan and Zheng, 2024).
A model can reach level 3
In a study of 1,203 people, those who rated advertising copy with the source hidden were most satisfied with copy written by ChatGPT-4 alone, level with copy that ChatGPT-4 finished from a human draft (Zhang and Gosline, 2023).
We found no test in which a model reached level 4 without a person's fact and a test on readers.