Using AI for Audience Research That Isn’t Persona Fan Fiction
Ask Claude or ChatGPT to “create a detailed persona for a UK B2B SaaS marketing manager” and you’ll get Sarah. Sarah is 34. Sarah is time-poor. Sarah values authenticity and struggles with proving ROI to her CMO. Sarah’s pain points are, in order: budget constraints, tool sprawl, and difficulty measuring attribution.
You already knew that. Everyone already knew that, because Sarah is a weighted average of ten thousand marketing blog posts about marketing managers, and those blog posts were themselves written by people summarising other blog posts. The model did exactly what it was built to do: it returned the most probable continuation of the phrase “B2B marketing persona.” The most probable answer is by definition the least differentiated one.
Here’s the shift that makes ai audience research actually useful: stop asking the model to generate and start asking it to read. You almost certainly have between 50,000 and 500,000 words of unstructured customer language sitting in systems nobody has ever run analysis against. Sales call transcripts. Support tickets. G2 and Trustpilot reviews. Sales email threads. Community posts. Win/loss notes. That corpus is the only audience research input that produces something your competitors can’t produce, because they don’t have it.
The input source is the entire argument. Everything else is prompt tinkering.
What “your own corpus” actually means in practice
Most small teams assume they don’t have enough data. They almost always do. Here’s what a five-person content team at a UK fintech I worked with turned up when they went looking, having initially told me they had “basically nothing”:
| Source | Volume found | Extraction route |
|---|---|---|
| Gong call recordings, 12 months | 412 calls, ~1.8M words | Gong API export to JSON |
| Intercom support conversations | 2,847 threads | CSV export, date-filtered |
| G2 + Capterra reviews (own + 4 competitors) | 690 reviews | Manual scrape, ~3 hours |
| Closed-lost notes in HubSpot | 233 deals | CRM report export |
| Sales team’s shared “objections” doc | 40 entries | Already text |
Nobody had ever looked at this together. The G2 competitor reviews alone are worth the afternoon: you are reading, in your buyers’ own words, the specific reasons people abandoned the tool they’re currently using. That’s a content calendar sitting in plain sight.
A caveat worth stating before you start. Call recordings and support tickets contain personal data, and UK GDPR applies whether or not you’re pushing it through an LLM. Two things keep this clean in practice: strip names, emails, phone numbers and company identifiers before anything leaves your systems (Microsoft Presidio is free and handles this reasonably well; a regex pass over obvious PII fields covers most of the rest), and use an API endpoint with zero data retention rather than a consumer chat interface. Anthropic’s API doesn’t train on your inputs by default; OpenAI’s API is the same, their consumer products aren’t necessarily. Check your DPA before you commit to a workflow, not after.
The extraction pass: three prompts that do the real work
The instinct is to dump everything in and ask for insights. Don’t. You get a competent summary that reads like every other competent summary. Instead, run narrow extraction passes where the model does one specific job per pass and you keep the intermediate output.
Pass one: verbatim mining. Feed the model batches of transcripts and ask it only to pull exact quotes matching a defined pattern. No paraphrasing, no synthesis.
For each transcript below, extract every verbatim sentence where the speaker
describes a workaround they currently use, a manual process they perform, or a
tool they use alongside ours. Return JSON:
{ "quote": "<exact text>", "call_id": "<id>", "speaker_role": "<role>",
"category": "workaround|manual_process|adjacent_tool" }
Do not paraphrase. Do not include quotes where the speaker is asking a
question rather than describing their situation. If a transcript contains
none, return an empty array for it.
Running this over those 412 calls at 100k-token batches cost roughly £61 on Claude Sonnet and produced 1,340 quotes. Precision was around 85% on a manual check of 60 sampled quotes, with the failures mostly being hypotheticals the model read as descriptions. That’s acceptable for research purposes and would have taken a human maybe 180 hours.
Pass two: clustering the verbatims. Now you have a quote corpus, cluster it. You can do this with the model, but embeddings are cheaper and less prone to inventing tidy categories. Run the quotes through text-embedding-3-large, dimension-reduce with UMAP, cluster with HDBSCAN, then have the model name each cluster from its 20 most central members. The whole thing is about 40 lines of Python and costs under £2 for a corpus this size.
Why bother with the statistical route? Because when you ask a model to categorise 1,340 quotes directly, it produces the categories it expects to find. Embeddings produce the categories that are actually there, including the ugly, specifically-shaped ones. The fintech team’s largest cluster, 187 quotes, was people describing a process of manually reconciling one report against another in Excel every Monday morning before a standing meeting. That’s not a persona trait. That’s a Tuesday-publishing opportunity with a known audience and known search behaviour.
Pass three: language capture. Separate pass, different question: what words do these people use that we don’t? Ask the model to list every noun phrase in the corpus describing the product category or a job to be done, ranked by frequency, and flag the ones absent from your own site copy.
The output from that fintech run, trimmed:
"month-end close" 341 occurrences ON SITE: no
"the reconciliation" 208 ON SITE: no
"getting the numbers to agree" 94 ON SITE: no
"financial close automation" 2 ON SITE: yes (47 pages)
"continuous accounting" 0 ON SITE: yes (12 pages)
They had built their entire content programme on two category terms that their customers essentially never said. Their buyers said “month-end close” and “reconciliation.” One of those phrases has UK monthly search volume in the low thousands; “continuous accounting” has, generously, double digits. Twelve pages of it.
The synthesis step, where the persona finally earns its place
Once you have clustered verbatims, you can ask the model for a persona and get something defensible, because every claim traces to quotes. The rule: no assertion without a citation into the corpus.
Using ONLY the clustered quotes provided, write an audience profile for the
segment represented by clusters 3, 7 and 11. Every claim must cite at least
two call_ids. Where you have fewer than two supporting quotes for something
you'd expect a profile to cover, write "INSUFFICIENT EVIDENCE" for that
element rather than inferring it. Do not use the words "pain point",
"journey", "seamless" or "values authenticity".
That INSUFFICIENT EVIDENCE instruction is the one that matters. You will get five or six of them, and they mark your actual research gaps: the questions to add to your next ten sales calls. A persona document honest enough to say “we don’t know how this person chooses between us and the incumbent” is worth more than one that confidently invents the answer.
You’ll also find the profile contradicts things you believed. The fintech’s deck said their buyer was the CFO. The corpus said the CFO signs, but the person with the problem, the vocabulary and the search behaviour is a financial controller three levels down who has never appeared in a single piece of their content. Reworking targeting around that finding is a planning decision rather than a research one, and the mechanics of translating corpus findings into a brief and a calendar belong to your AI-assisted content strategy and planning workflow.
What this costs and how long it takes
A first full run on a corpus of 1-2 million words: two days of your time, mostly export wrangling and PII stripping, plus £70-£120 in API spend. The ongoing version is cheaper. Set the verbatim-mining pass to run monthly against new calls and tickets, append to the quote corpus, re-cluster quarterly. That’s maybe £15 a month and an hour of attention.
Compare that to commissioning qualitative research: £8,000-£25,000 for a UK agency to run 12-15 depth interviews, six weeks of turnaround, and a sample size two orders of magnitude smaller than your call archive. The agency work is better in one specific way, which is that a good interviewer asks the follow-up question nobody thought to ask. Your corpus can’t do that. It can do volume, recency and continuity, and it costs roughly what a single sponsored LinkedIn post costs.
Where it goes wrong
Three failure modes show up repeatedly.
Sales calls over-represent people who took a sales call. If you only mine Gong, you learn about the buyers who reached the demo stage and nothing about the ones who read two blog posts and left. Balance with support tickets (existing customers, real usage) and competitor reviews (people who chose someone else). Three sources minimum.
Recency matters more than you’d think in fast-moving categories. Weight the last six months heavily, or at least tag by date and check whether your biggest cluster is mostly quotes from 2024. A theme that was dominant eighteen months ago and has since decayed will still dominate an unweighted corpus.
And the temptation to let the model fill gaps will return, quietly, at the point where the profile looks thin and you have a brief due. When you catch yourself prompting “based on this, what else might this audience care about?”, you have re-entered fan fiction. The answer is another extraction pass or ten more calls, not a better prompt.
Start with the G2 reviews for your three closest competitors. It takes an afternoon, needs no engineering, no PII handling and no budget approval, and by the end of it you will have between 30 and 80 verbatim complaints about the products your buyers are currently using. Put them in a spreadsheet, cluster them, and see how many map to something you’ve published.