Draft, Then Interrogate: Getting Past the First Output
The first thing a model gives you is not a draft. It’s a straw man: a plausible-sounding, statistically average arrangement of what everyone else has already published on the topic. Treat it as a draft and you’ll spend forty minutes tidying a document that was never worth tidying. Treat it as an opponent’s opening statement, and something useful happens.
This is the single biggest difference I see between content teams who get real value out of an ai writing workflow and teams who quietly stop using AI after three months because “everything came out the same.” The prompt isn’t the problem. The editing isn’t the problem. The missing step is the interrogation pass that sits between them.
Why the first output is always average
Worth being precise about the mechanism, because it changes how you respond to it.
A language model generating a paragraph about, say, B2B lead scoring is producing the most probable continuation given everything it has seen. Most of what it has seen about B2B lead scoring is mediocre agency blog content. So the output regresses to that mean: hedged, structurally symmetrical, full of “it’s important to note” and three-item lists where two items would do. It isn’t lying. It’s averaging.
You cannot prompt your way out of averaging entirely. You can ask for “a contrarian take” and get the most probable contrarian take, which is its own genre of sludge. What you can do is treat the average as a starting position and then apply pressure to it, the way a decent editor applies pressure to a junior writer’s first go.
The interrogation pass, concretely
Here’s the workflow I run, and the rough time cost of each stage for a 1,200-word piece:
| Stage | Time | What you’re actually doing |
|---|---|---|
| Brief + first output | 8 min | Generating the straw man |
| Claim audit | 12 min | Marking every unsupported assertion |
| Counter-argument pass | 10 min | Forcing the model to attack its own draft |
| Specificity forcing | 15 min | Replacing categories with instances |
| Human rewrite | 35 min | Writing the thing properly |
Total: about 80 minutes. Compare that to roughly 110 minutes for the same piece written from scratch, or 95 minutes for “generate then edit” where the edit is polishing rather than interrogating. The saving is modest. The quality difference is not.
Stage one: the claim audit
Take the draft back into the model and ask it to turn on itself. The prompt I use, near enough verbatim:
Below is a draft you produced. Go through it claim by claim.
For each factual or causal assertion, output a row:
CLAIM | EVIDENCE TYPE | CONFIDENCE | WHAT WOULD DISPROVE IT
Evidence type is one of: cited study, industry-common-knowledge,
plausible-but-unverified, invented.
Be harsh. If you cannot name a source, say plausible-but-unverified.
Do not defend the draft.
Claude Opus 5 and GPT-5 both handle this well; Gemini 2.5 Pro tends to be a little too generous with itself unless you repeat “be harsh” in the final line. What comes back is genuinely uncomfortable. On a recent draft about email re-engagement, 14 of 19 claims came back as plausible-but-unverified. Three were flatly invented, including a confident “studies show a 23% average uplift” that traced to nothing at all.
That table is your research list. It is also, usefully, a map of where the piece is hollow. A section where every row says plausible-but-unverified is a section you should probably cut rather than shore up.
Stage two: demand the counter-argument
Now make it argue the opposite. Not “give me a balanced view”, which produces mush, but a specific adversarial brief:
You are a sceptical practitioner who disagrees with this draft's
central argument. You have 15 years' experience and you think this
piece is naive. Write 300 words explaining why it's wrong, using
specific scenarios where following its advice would cost money.
Two things come out of this. First, the strongest objection, which you now have to answer in the piece rather than pretend doesn’t exist. Second (more often than you’d expect) the realisation that the objection is better than your original argument and the piece should be rebuilt around it.
I ran this on a draft arguing that small teams should build a content calendar twelve weeks ahead. The counter-argument came back pointing out that for a three-person team, a twelve-week calendar is mostly a record of decisions made before you had the data to make them, and that the real constraint isn’t planning horizon but the two-week gap between publishing and having enough traffic data to know if it worked. That became the piece. The original draft went in the bin, which is the correct outcome for a straw man.
Stage three: force specificity
This is where most of the genericness dies. The model writes in categories: “businesses”, “tools”, “significant improvements”. Your job is to convert every category into an instance.
The mechanical version:
Rewrite this section. Every noun phrase that names a category must
be replaced with a named example. Every quantity word (many, most,
significant, substantial) must be replaced with a number or removed.
If you do not know a number, write [NUMBER NEEDED: what exactly].
Do not invent figures.
“Many teams find their content workflow tools don’t integrate well” becomes “Notion and Ahrefs don’t talk to each other, so your brief lives in one place and your keyword data in another, and someone copies between them.” The second version might be wrong for a given reader. The first version is unfalsifiable, which is worse.
Those [NUMBER NEEDED] markers are the useful artefact. On a 1,200-word piece you’ll typically get 6 to 10 of them. Each one is a fifteen-minute research task or a sentence you delete. Both are fine. What isn’t fine is leaving the hedge in.
What this looks like against real numbers
An agency content lead I worked with in Bristol ran this across 22 pieces over one quarter. Her before-and-after, measured on pieces of comparable length and topic difficulty:
- Average time to publishable draft: 2h 40m down to 1h 50m
- Pieces requiring a second full rewrite after client review: 7 of 22 in the previous quarter, 1 of 22 after
- Claims removed at the audit stage: 94 across the 22 pieces, of which 11 were fabricated specifics
That last figure is the one worth sitting with. Eleven invented facts that would have gone out under a client’s name, caught by a pass that costs twelve minutes. The productivity argument for interrogation is decent. The not-publishing-nonsense argument is overwhelming.
Where it fits in your existing process
Slot the interrogation between generation and editing, not instead of either. You still need a proper brief going in, and you absolutely still need a human writing pass at the end, because the model cannot do voice and the interrogated draft will read like a well-sourced robot. The broader question of how drafting, voice and editing chain together is worth reading up on separately; there’s a fuller treatment in our guide to drafting, brand voice and editing workflows if you want the surrounding process rather than just this stage.
One practical warning. Interrogation works in a fresh context more reliably than in a long thread. If you’ve been chatting with the model for twenty turns about how good the draft is, it has absorbed that framing and will defend the work. Paste the draft into a clean conversation with no history. Claude Projects and ChatGPT’s custom GPTs both let you save the three prompts above as a standing setup, which removes the temptation to soften them.
The uncomfortable bit
Some of your pieces will not survive this. You’ll run the counter-argument pass and discover the article has nothing to say, at which point the honest move is to not publish it. That’s a feature. A content programme that ships fourteen pieces a month, nine of which are generic, is worse off than one shipping six that aren’t, and the interrogation pass is a cheap way to find out which is which before you’ve spent the writing time.
Start with your next piece. Run the claim audit only, skip the other two stages, and see how many rows come back as plausible-but-unverified. Whatever that number is, it was always there. You just hadn’t asked.