Feeding AI Your Own Research So Drafts Aren’t Hollow
The complaint arrives in the same shape every time. Someone on a two-person content team runs a brief through Claude or ChatGPT, gets back 900 words that are grammatical, on-topic, structurally sound, and completely empty. No specifics. No numbers anyone could check. A paragraph about “the importance of understanding your audience” that could sit on any of four thousand websites. The reflex is to blame the model, then to blame the prompt, then to spend two hours rewriting until it sounds like a human wrote it badly rather than a machine wrote it well.
That diagnosis is wrong, and the wrongness is expensive. Hollow output is not a generation problem. It is an input problem. A language model asked to write about customer onboarding with nothing but the phrase “customer onboarding” in front of it will produce the statistical average of everything ever written about customer onboarding, because that is the only material available. The average is hollow by construction. Give the same model a 40-minute transcript of your head of customer success describing why three accounts churned in Q2, and the drafting step stops being invention. It becomes assembly.
That distinction is the whole of a working ai content creation workflow. Everything below is about the part that happens before anyone writes a word.
What “primary material” actually means
Four categories, and most teams already have three of them sitting unused.
Interviews. Not just SME interviews for a specific piece. Sales calls, customer success check-ins, onboarding sessions, the internal retro where someone explained why the migration took eleven weeks instead of four. Gong, Grain and Fathom all export transcripts. If you use Google Meet, transcripts are in Drive already.
Proprietary data. Your product analytics, survey results, benchmark numbers, support ticket volumes. A single honest number nobody else has (“median time-to-first-value across our 340 accounts is 9 days, but the top quartile hits it in 2”) is worth more than 400 words of framework.
Internal documents. Positioning docs, battlecards, the objection-handling wiki page, Slack threads where someone argued about pricing. Your sales enablement library is a compressed record of what customers actually resist.
Customer language. Review sites, support tickets, the free-text field on your NPS survey, community posts. This is where you learn that nobody says “workflow automation” but everyone says “stop doing the same thing every Monday.”
A worked example, start to finish
Take a brief: “Why finance teams stall on spend management rollouts” for a mid-market fintech, target 1,400 words.
The hollow version. Prompt the model with the title, a note about audience, and a request for eight sections. Time spent: four minutes. Output: a piece that mentions “change management,” “stakeholder buy-in,” and “phased implementation.” Editing time before it is publishable: about 3 hours, because every section needs its generic claim replaced with something true.
The fed version. Before drafting, spend 90 minutes assembling context:
- Two customer interview transcripts from calls the CS team already recorded (roughly 9,000 words combined)
- An export of 60 support tickets tagged
rolloutfrom the last six months - The internal post-mortem doc on the Brightwell account, where rollout took 5 months
- Win/loss notes from 12 closed-lost deals in the segment
That is maybe 22,000 words of raw material, comfortably inside a 200k context window. Then the drafting prompt does something different: it asks for assembly, not invention.
You have four sources below: two customer interview transcripts,
60 support tickets, one internal post-mortem, and win/loss notes.
Before drafting, do this:
1. List every distinct reason a rollout stalled, with the source
and a direct quote for each. Do not generalise.
2. Flag any reason that appears in 3+ sources.
3. Note any claim I might want to make that the sources do NOT
support.
Then stop. I will review before you draft.
That stop instruction matters more than anything else in the prompt. What comes back on a real run of this looks roughly like:
Reasons appearing in 3+ sources:
- Approval chain not mapped before go-live (7 tickets, both
interviews, post-mortem p.3)
- Finance owns the tool, ops owns the process, nobody owns the
rollout (both interviews, 4 win/loss notes)
- Existing spreadsheet workflow runs in parallel for months
"just in case" (11 tickets, 1 interview)
Single-source, still interesting:
- "We didn't tell anyone it was mandatory until month three."
(Brightwell post-mortem, p.5)
NOT supported by your sources:
- Any claim about cost of delay. No source quantifies it.
- Any claim that training quality is a factor. Two interviews
explicitly say training was fine.
Look at what that last block does. It stops you writing the paragraph about training that every competitor has written, and it stops you inventing a cost-of-delay statistic. The “not supported” instruction is the single highest-value line in the whole prompt, and almost nobody includes it.
Draft from that, and the model writes about parallel spreadsheet workflows and unmapped approval chains, because those are the things in front of it. Editing time drops to about 45 minutes, and the piece contains three claims no competitor can make.
Net: 90 minutes of prep plus 45 of editing beats 4 minutes of prompting plus 3 hours of rescue. You are also 15 minutes ahead on the clock, and the output is a different category of thing.
Building the context layer so prep isn’t 90 minutes every time
The obvious objection: 90 minutes per piece does not scale on a three-person team publishing eight pieces a month. Correct. The prep cost drops sharply once material is organised rather than hunted.
Three approaches, in ascending order of effort.
Project files. Claude Projects and ChatGPT Projects both let you attach documents that persist across conversations. Load your positioning doc, brand guidelines, three best-performing published pieces, and a customer language glossary once. Every conversation in that project inherits them. Setup: an afternoon. Practical ceiling: roughly 20 to 30 documents before retrieval gets loose.
A research repo. A folder per topic cluster, containing transcripts, data exports, and a running notes file. Notion works, Obsidian works, a Google Drive folder works. The discipline is that every customer call, every survey, every data pull gets filed to a cluster when it happens, not when you need it. A content lead who does this consistently has 4 to 6 usable sources per brief waiting before the brief is written.
A quote bank. One spreadsheet. Columns: quote, source, date, topic tag, permission status. Every time anyone on the team hears something good on a call, it goes in. After six months you have 200 lines, and the “find me a customer voice for this section” problem is a filter, not a research project. This is the highest return per unit of effort of anything in this piece.
Tooling beyond that gets heavier: NotebookLM will take up to 50 sources and answer against them with citations back to the source document, which makes it genuinely good for the extraction step. Gemini’s integration with Google Drive means a research folder can be queried directly. Neither is necessary. The spreadsheet does most of the work.
Loading order matters more than you’d think
Context is not a bucket you pour things into. Position affects what surfaces.
Put raw material first and instructions last. Models attend more reliably to instructions at the end of a long context than buried at the top, and a 20,000-word source dump will swallow a 200-word brief placed before it. Structure it: sources, then the brief, then the task.
Label each source explicitly. <interview_1 speaker="Head of Finance, 400-seat SaaS" date="2026-07-14"> beats an unlabelled wall of text, because the model can then attribute claims correctly and you can spot-check them. When it writes “one finance lead described approvals as ‘a phone tree with no operator’,” you know exactly where to verify.
Do not pre-summarise. The instinct to condense 9,000 words of transcript to 500 words of bullets before feeding it in destroys the thing you came for. Specificity lives in the phrasing, the hesitation, the odd analogy someone reaches for. Summarising strips all three and hands the model back the generic input you were trying to escape. If a source genuinely will not fit, cut whole sections rather than compressing everything.
Where this connects to drafting and editing
Fed context changes what the drafting stage is for. When the model has 22,000 words of primary material, prompting for “engaging, conversational tone” is close to pointless: tone follows from having something to say. The real drafting work becomes structural (what order do these findings go in) and voice-level (does this sound like us). That is a separate discipline with its own mechanics, covered in our guide to drafting, brand voice and editing workflows, and it gets considerably easier when the raw material is already in the room.
One rule bridges the two stages: every specific claim in the draft must trace to a source you loaded. Run a verification pass where you ask the model to list each factual claim with its source, then check the three or four that matter most. Models will still occasionally attribute a quote to the wrong interview, or smooth a number. On a piece with 22,000 words of context you might find one error per draft. Worth eight minutes to catch.
What this costs, honestly
For a team publishing eight pieces a month, expect the first two months to feel slower. Building the quote bank and filing habit is overhead before it is leverage. Somewhere around month three, prep per piece drops from 90 minutes to 25, because the material is already there and organised.
There are briefs where you have nothing. New market, no customers yet, no data. In those cases the honest move is to go get primary material (three 30-minute interviews with people in the segment) or accept that the piece will be thinner and treat it accordingly. Feeding the model a competitor’s blog post as “research” produces a laundered version of their thinking, which is worse than generic because it is confidently someone else’s.
Start with the quote bank. One spreadsheet, five columns, and a standing instruction that anyone who hears something good on a call writes it down. Six weeks from now the next brief will have somewhere to begin.