AI Content Marketing
§2 Section 2 of 6 2,649 words · 12 min

Drafting, Brand Voice and Editing Workflows

Most content teams that “use AI” are really doing one thing: pasting a brief into a chat window, getting 900 words back, then spending 90 minutes unpicking it. The time saved on the first draft gets eaten by the rewrite, and the piece still reads like it was assembled by a committee of LinkedIn thought leaders. That’s not an AI content creation workflow. That’s a slot machine.

A real workflow has fixed inputs, defined stages, a named owner at each handoff, and a quality gate that fails loudly. This page covers the production half of that: how to get a usable draft, how to make it sound like your brand rather than everyone’s brand, and how to edit efficiently when the raw material has different failure modes than a human first draft. Everything here assumes a team of one to five people with no engineering support and a budget that tops out around £200 a month in tooling.

The Four-Stage Draft Pipeline

The single biggest improvement you can make is refusing to generate a draft in one shot. Split it.

Stage 1: Evidence assembly (human, 30-45 minutes). Before any model sees the brief, you collect the raw material that makes the piece non-generic. For a B2B SaaS piece on invoice financing, that means: two customer quotes from your Gong or Fathom call recordings, the actual pricing of three competitors, one internal data point (churn, conversion, support ticket volume), and two links to primary sources. Put all of it in a single document. If you can’t assemble this, the piece will be generic no matter what model you use, because there’s nothing in it that isn’t already in the training data.

Stage 2: Outline generation (AI, 10 minutes). Feed the evidence pack plus the brief, ask for three structurally different outlines, and pick one. Not “write an outline” but “give me three: one problem-first, one contrarian, one framework-led. For each, tell me which section carries the most original value.” The third clause is the useful one. It forces the model to rank sections, which surfaces the parts that are just connective tissue.

Stage 3: Section-by-section drafting (AI, 20-40 minutes). Draft one H2 at a time, 300-500 words per call, passing the outline and the previously written sections as context. This sounds slower. It isn’t, because full-article generation produces uniform paragraph rhythm and a narrowing vocabulary as the context window fills, and fixing that costs more than the extra prompts. Section drafting also lets you kill a bad section at 400 words rather than at 2,000.

Stage 4: Structural edit (human, 45-60 minutes). Not line editing. You’re checking whether the argument holds, whether the evidence actually supports the claims, and whether any section could be deleted without loss. Line editing comes after, and it’s faster once the structure is right.

Total: roughly two and a half hours for a 2,000-word piece that would take a good writer four to six. The saving is real but it’s 40-50%, not the 10x the tooling vendors imply. Teams that report 10x are either measuring only Stage 3 or shipping work that isn’t good.

Prompting for Sections That Don’t Sound Like Everything Else

Three patterns do most of the heavy lifting.

Negative constraints beat positive ones. “Write in a conversational tone” produces the house style of every model, which is a sort of chipper explainer voice. “Do not use the words leverage, seamless, robust, landscape, or delve. Do not begin any paragraph with ‘In today’s.’ Do not use the construction ‘It’s not just X, it’s Y.’ Do not end sections with a summary sentence” produces something noticeably different. Keep a running banned list in your prompt template and add to it every time a tic annoys you. Mine is currently 34 items long.

Give it something to react to. Models are much better at critique and revision than at cold generation. Instead of “write a section on why attribution models fail for content,” try: “Here’s the standard argument for last-touch attribution [paste 200 words]. Write a section explaining where this breaks for content marketing specifically, using the evidence in the pack. Disagree with the standard argument where the evidence supports disagreement.”

Specify the unit of thought, not just the word count. “Write 400 words” gets you four paragraphs of similar length and similar weight. “Write this section as: one claim, one piece of evidence from the pack, one objection to the claim, one response to the objection. Roughly 400 words, but let the structure set the length” gets you something with shape.

A worked example. Brief: a piece for a fintech client about why SME lending decisions take so long. Weak prompt: “Write a 500-word section about the causes of slow SME lending decisions.” Output: manual processes, legacy systems, regulatory requirements, risk aversion. All true, all in every other article on the subject.

Better prompt: “Section on why SME lending decisions take 4-6 weeks. Use these three data points from our evidence pack: [pack]. Structure: lead with the specific bottleneck our ops team identified (document chasing, 11 days average), then explain why the obvious fix (open banking data pulls) only removes about 3 of those days, then what actually removes the rest. No mention of ‘legacy systems’ or ‘digital transformation.’ Write for a credit ops manager who already knows the process, so don’t explain what underwriting is.”

The second prompt works because it contains information the model doesn’t have, an argument structure, a banned phrase list, and a reader specification. Three of those four come from Stage 1.

Brand Voice: Why “Professional But Friendly” Does Nothing

Ask a model to write in a “professional but friendly” voice and you get the model’s default, because that string describes roughly 60% of published business English. The instruction carries no information.

Voice instructions only work when they’re falsifiable. “We use sentences of varying length, with at least one sentence under six words per every 150 words” is checkable. “We never use rhetorical questions as section openers” is checkable. “We write in second person singular, addressing one reader, and we never say ‘businesses’ when we mean ‘you’” is checkable. “Approachable yet authoritative” is not.

The second problem is that voice guides written for humans assume shared context. A human writer reads “confident, not arrogant” and applies twenty years of reading. A model reads it and produces nothing in particular. You need a different artefact: rules with examples, stated as pairs.

RuleWe writeWe don’t write
Name the actor“Your finance team approves it”“Approval is obtained”
Numbers not adjectives“Cuts review time from 9 days to 4”“Dramatically reduces review time”
No hedging stacks“This usually works”“This may potentially help in some cases”
Concrete nouns“Xero, Sage, QuickBooks”“Your accounting software stack”

Four rules with examples outperform two pages of adjectives. I’ve tested this on the same brief with the same model: a 300-word rules-and-examples guide produced drafts requiring roughly half the line editing of drafts produced from a 1,200-word traditional brand voice document.

Building that artefact properly is its own piece of work, and there’s a full method for it in our guide to writing a brand voice guide an LLM can actually use, including how to extract rules from your existing best-performing content rather than inventing them.

One practical note on storage: put the voice guide somewhere the model reads on every call, not somewhere you paste it when you remember. In ChatGPT that’s a Project with the guide in project instructions. In Claude that’s a Project with the guide in the knowledge base. In Gemini that’s a Gem. If your team is using raw chat windows with no persistent context, your voice consistency depends entirely on whoever is prompting that day, which is the thing you’re trying to fix.

Voice Calibration: Getting From 70% to 90%

Even a good voice guide gets you to roughly 70% on the first try. Closing the gap is an iteration loop, and it takes about two hours once.

Take three published pieces you’re happy with. Paste one into the model alongside your voice guide and ask: “Here’s our voice guide and here’s a piece that follows it well. What’s in the piece that isn’t captured in the guide?” The answers are usually specific and useful: sentence-opening variety, a habit of putting the caveat before the claim, a preference for colons over dashes, the way you handle numbers in running text.

Then reverse it. Generate a draft, put it next to the human-written piece, and ask the model to list every stylistic difference. You’re mining for the tells. Common ones across models: three-item lists everywhere (models love a tricolon), every paragraph roughly the same length, heavy use of “not just X but Y,” a compulsion to end sections with a forward-looking sentence, and transitional adverbs at the start of sentences (“Moreover,” “Additionally,” “Furthermore”).

Add what you find to the guide as new rules. Two rounds of this typically gets you to the point where a light line edit is enough. A useful checkpoint: pull five paragraphs from a finished AI-assisted piece and five from something a team member wrote, shuffle them, and ask a colleague who wasn’t involved to identify which are which. If they can do better than chance on more than seven out of ten, the voice work isn’t done.

Editing AI Drafts: Different Failure Modes, Different Passes

Human first drafts fail by being under-structured, repetitive, or unfinished. AI drafts fail differently, and applying your normal editing process to them wastes time. Run three separate passes.

Pass one: fact and specificity audit (15 minutes). Go through with a highlighter for every claim that has a number, a date, a name, or a citation. Verify each one. Models fabricate statistics with total confidence, and they fabricate them plausibly, which is worse. They also invent attributions: a statistic that genuinely exists gets credited to the wrong source. In my experience roughly one in six numerical claims in an unverified draft is wrong, misattributed, or unfindable. Also flag hollow specificity: “studies show,” “many companies,” “industry leaders.” Each one is either replaced with a real source or cut.

Pass two: compression (20 minutes). AI drafts are typically 20-30% longer than they need to be. The fat is predictable: sentences that restate the previous sentence in different words, paragraphs that open with a throat-clearing framing sentence, section closers that summarise what you just read. Cut the first sentence of most paragraphs and see if anything’s lost. Usually not. Delete every sentence beginning “By doing this,” “This means that,” or “Ultimately.” Target a 20% word reduction and you’ll find it comfortably.

Pass three: voice and rhythm (20 minutes). Read aloud, or use text-to-speech at 1.5x. Rhythm problems are inaudible on screen and obvious in the ear. You’re listening for uniform sentence length, repeated openers, and the flat delivery that comes from every paragraph having exactly the same structure. Fix by varying: break a long sentence, merge two short ones, start one paragraph with a fragment.

Doing these as three passes rather than one combined edit is faster, because each pass has a single question. Combined editing means you’re context-switching between fact-checking and rhythm, and you’ll miss things in both.

Tooling: What’s Actually Worth Paying For

For a three-person team, a working stack looks like this.

Drafting model. Claude or ChatGPT at £15-17 per seat per month. Both are good. Claude tends to follow long constraint lists more reliably in my testing, ChatGPT’s Projects and custom GPTs are easier to share across a team. If you’re picking one, pick based on which interface your team will actually use daily rather than benchmark scores.

Persistent context. Use the native Projects feature rather than a third-party wrapper. Set up one Project per content pillar, with the voice guide, three exemplar pieces, and the ICP definition in the knowledge base. Cost: nothing extra.

Transcript mining. Fathom (free tier is genuinely usable) or Fireflies for call recordings. This is where Stage 1 evidence comes from. Customer calls are the single richest source of non-generic material most teams already have and never use.

Fact-checking assist. Perplexity Pro at £17/month, used specifically to verify claims in Pass One. Its citation behaviour is better suited to verification than general-purpose chat.

Editing. Human. There is no tool that does Pass Two well, because compression requires knowing what the piece is for.

Total: about £70-100 a month for a small team. What I’d skip: dedicated “AI writing platforms” at £80-200/month that wrap a frontier model with templates. You’re paying a premium for prompt templates you could write better yourself in an afternoon, and you lose the ability to move between models.

Governance That Fits on One Page

Three rules, written down, enforced at the point of publishing.

Rule one: named owner. Every piece has one human whose name is on it internally, who is accountable for every factual claim. Not “the AI got it wrong.” The owner got it wrong.

Rule two: disclosure position. Decide once whether AI-assisted content carries a disclosure, and where. Most UK B2B teams land on no disclosure for standard content and explicit disclosure for anything where authority matters (research reports, expert commentary, regulated content). Financial services and health teams have less latitude here: check with compliance before you set the policy, not after.

Rule three: the no-AI list. Some things don’t go through this workflow. Founder opinion pieces, customer case studies where the customer’s words matter, crisis communications, anything with legal exposure. Write the list, put it in the content calendar template as a checkbox.

Add a fourth if you work with freelancers: state your position on AI use in the brief. “AI-assisted drafting is fine, AI-generated final copy is not, and you’re accountable for accuracy” is a reasonable position that most good freelancers will accept. Silence on this gets you drafts you’re paying human rates for and receiving model output.

What to Measure

Track two things for the first quarter, and resist adding more.

Time per published piece, split by stage. Log it in a spreadsheet for eight weeks. You need to know whether the workflow is actually saving time or just moving it from drafting to editing. If Stage 4 is taking longer than 60 minutes on a 2,000-word piece, your Stage 1 evidence pack is too thin or your voice guide isn’t specific enough. That’s diagnostic, and it tells you exactly where to invest.

Quality parity, not quality improvement. Compare engagement metrics on AI-assisted pieces against your baseline from the previous two quarters. Scroll depth, time on page, and conversion to whatever your next step is. You’re looking for “no worse,” because the case for this workflow is speed at equal quality, not better content. If AI-assisted pieces underperform by more than 15% on scroll depth, something in the pipeline is producing work that loses readers partway through, and that’s almost always a specificity problem in the middle sections.

A note on search: as of late 2026, Google’s position is that AI assistance is fine and low-quality content is not, which is the same position they’ve held about content produced any other way. The teams getting hit are the ones publishing volume without evidence packs. The workflow above is, incidentally, quite a good defence, because evidence assembly is the thing that produces the original information Google’s systems are looking for.

Start with one pillar topic. Run four pieces through the full four stages, log the time, and compare the output to what you were shipping in June. If the fourth piece isn’t noticeably better than the first, the problem is almost certainly Stage 1, and the fix is boring: better evidence, gathered by a human, before anything gets generated.

In this section

The supporting pages under this subject.