AI Content Marketing

AI-Assisted Pillar and Cluster Mapping at Speed

Most content teams build their site architecture around keyword similarity because that’s what the tools make easy. Feed 4,000 keywords into a clustering tool, set a similarity threshold, get 180 groups back. It looks like strategy. It’s actually a very expensive way to reorganise a keyword export.

The problem is that keyword similarity and buyer intent are only loosely correlated. “Payroll software UK”, “best payroll software”, “payroll software pricing” and “payroll software for small business” will cluster together in every SERP-overlap or embedding-similarity tool you can name. But those four queries come from three different people at three different points in a buying decision, and a single 2,800-word page trying to serve all of them will serve none of them well. Meanwhile “how do I handle a mid-year payroll switch” ends up in a different cluster entirely, despite being the single highest-intent question your sales team hears in month two of every deal.

So the reframe is this: cluster by buyer question and decision stage first, then let keyword data tell you how to title and structure what you’ve already decided to build. AI is genuinely good at that sorting job, and it’s the part that used to take three days of whiteboarding.

Why keyword-similarity clustering produces commercially useless maps

Here’s a real pattern. A B2B fintech client ran a standard clustering pass on 3,200 keywords using Keyword Insights (SERP-overlap method, 3 URL match threshold). Output: 214 clusters. They built 40 pages over seven months against the top clusters by volume.

Result after seven months: traffic up 71%, demo requests from organic up 4%. Classic. Every cluster they’d prioritised was an awareness-stage definitional query, because awareness queries have the highest search volume and therefore sorted to the top. They’d built an encyclopaedia. Nobody buys from an encyclopaedia.

The structural failure has three parts:

Volume sorting is stage sorting in disguise. Search volume is roughly inversely proportional to decision-stage proximity. “What is reverse factoring” gets 1,900 UK searches a month. “Reverse factoring vs invoice discounting for construction subcontractors” gets 20. The second one closes deals. Sort by volume and you systematically deprioritise revenue.

Similarity thresholds split real decisions and merge fake ones. The decision “do we switch providers mid-contract” spans comparison queries, migration queries, contract queries and cost queries. No similarity algorithm groups those, because they share almost no lexical or SERP overlap. Yet all four belong on one page, or one tight set of three.

The map has no internal logic for a model to follow. This matters more every quarter. When an LLM assembles an answer about your category, it’s doing entity and relationship resolution. A site where 40 pages all define adjacent variations of the same concept gives a model no signal about which page is authoritative on what. A site where each page answers one distinct buyer question, with explicit links up to the pillar and sideways to sibling decisions, is navigable.

The method: question-first clustering with AI doing the sort

This takes about four hours end to end for a mid-sized programme. I’ll use the payroll software example throughout because it’s boring enough to be clear.

Step 1: harvest actual questions, not keywords (45 minutes)

You need 200 to 400 real buyer questions. Keyword tools give you queries; you want questions. Sources, in rough order of value:

  • Sales call transcripts. Gong, Fathom or tl;dv exports. Pull 15 to 20 calls.
  • Support tickets and chat logs, last 90 days.
  • Won/lost deal notes from your CRM.
  • Reddit and industry forum threads (r/AskUK, r/smallbusinessuk, AccountingWEB for this example).
  • Your own search console queries filtered to question modifiers.

Dump the raw text into Claude or ChatGPT in batches. The prompt that works:

Below are transcripts from 18 sales calls for a UK payroll software
product. Extract every distinct question a prospect asked or implied,
verbatim where possible. Do not summarise or group them. Do not
deduplicate near-identical phrasings from different calls — I want
frequency signal. Output as a numbered list.

That “do not group” instruction matters. Both models want to helpfully tidy at this stage, and tidying destroys the raw material you need for the next step. From 18 calls you’ll typically get 130 to 190 questions.

Step 2: classify by decision stage and question type (30 minutes)

Now you do group, but along two axes at once. Feed the question list back with a rubric you define, not a generic funnel:

Classify each question against two dimensions.

DECISION STAGE:
- Problem-aware (knows something is wrong, doesn't know the category)
- Category-aware (evaluating whether this type of solution fits)
- Solution-comparing (comparing named options)
- Implementation-anxious (has decided, worried about the switch)
- Post-purchase (already a customer, needs to do a thing)

QUESTION TYPE:
- Definitional / Cost / Risk / Process / Comparison / Compliance / Proof

Output a table: question | stage | type | my confidence (high/med/low).
Flag anything you'd classify differently with more context.

The two-axis grid is where the commercial map appears. A representative slice:

QuestionStageType
Can I run payroll myself or do I need a bureau?Category-awareDefinitional
What happens to our RTI submissions if we switch in October?Implementation-anxiousRisk
Is Xero payroll enough or do we need something dedicated?Solution-comparingComparison
How much does payroll software cost for 40 staff?Category-awareCost
Who’s liable if the software files a wrong FPS?Implementation-anxiousCompliance
How do I correct a submission after payday?Post-purchaseProcess

Ask for the confidence column and act on it. On a 160-question run you’ll get 20 to 30 low-confidence rows, and those are almost always the interesting ones: questions that genuinely sit at a stage boundary, which usually means they deserve their own page rather than being absorbed.

Step 3: build clusters from the grid, not from the keywords (60 minutes)

Count your grid. A typical distribution from this exercise:

                    Problem  Category  Comparing  Impl.  Post
Definitional            18        22          3      1      2
Cost                     4        19         14      6      1
Risk                     9         7          8     24      3
Process                  2         5          2     11     21
Comparison               1         8         27      4      0
Compliance               6        11          5     17      8
Proof                    0         6         19      3      1

Two things jump out of that 160-question set. Implementation-anxious risk and process questions total 35, and almost nobody publishes for them because the keyword volume is negligible. And the comparison column holds 27 direct comparison questions against 8 category-aware ones, which tells you the comparison content needs to be named-competitor specific, not “how to choose”.

Now cluster. One cluster equals one buyer decision, which typically absorbs three to nine questions. The prompt:

Group these classified questions into clusters where each cluster
represents ONE decision a buyer is trying to make. A cluster may span
question types if they serve the same decision. Target 4-8 questions
per cluster. For each cluster give me: the decision in the buyer's own
words, the questions inside it, and the single page that would answer it.
Do not create clusters based on shared keywords.

From 160 questions you’ll land around 24 to 30 clusters. Compare that to the 214 from the similarity tool. That’s the whole point: you’ve gone from 214 topically tidy groups to 27 decisions, which is roughly how many decisions actually exist in a considered B2B purchase.

Step 4: map clusters to pillars by stage weight (45 minutes)

Pillars are not “big topics”. A pillar is the page that owns a stage of the decision and routes to the clusters beneath it. For the payroll example, four pillars fell out:

  1. Payroll software for UK small businesses (category-aware, 9 clusters beneath)
  2. Choosing between payroll options (solution-comparing, 7 clusters)
  3. Switching payroll providers (implementation-anxious, 8 clusters)
  4. Running payroll: tasks and corrections (post-purchase, 3 clusters)

Pillar 3 did not exist in the similarity-based map at all. Its 8 clusters had been scattered across 14 different keyword groups. It became the highest-converting section of the site inside four months, at a fraction of the traffic: 2,100 sessions a month against the category pillar’s 14,000, but 31 demo requests against 6.

This pillar-and-cluster layer is where most of the leverage in AI-assisted content strategy and planning actually sits, because every downstream decision (briefs, internal links, refresh priority) inherits its logic from this map.

Step 5: validate against keyword data, in that order

Only now do you open Ahrefs or Semrush. For each cluster, you’re answering three questions: does search volume exist for this decision, what phrasing do people use, and is anyone else serving it. Keep a cluster with zero search volume if your sales team hears the question weekly. That’s a page that will earn links, get cited in LLM answers and close deals from direct traffic, and it will never appear in a keyword tool.

One practical filter: if a cluster has fewer than 40 combined monthly searches AND sales has never heard the question, drop it. Everything else stays.

Making the map navigable for models

The internal linking pattern matters as much as the pages. Three rules, applied mechanically:

Every cluster page links up to its pillar with anchor text naming the decision, not the keyword. “Switching payroll providers mid-year” beats “payroll software”. Every cluster page links sideways to the two or three sibling clusters that represent adjacent decisions at the same stage, and states why in the sentence: “before you can work out the migration timing, you need to know whether your current provider charges an exit fee.” Pillars link down to every cluster beneath them with a one-line description of the decision each answers.

That gives you an explicit, machine-readable decision graph. When a model is assembling an answer to “should I switch payroll providers in October”, the relationships are legible on the page rather than inferred from lexical overlap. I’ve watched clusters built this way get cited in ChatGPT and Perplexity answers within five weeks of publishing, at page-level rather than site-level, which almost never happens with definitional content.

Where this breaks

Two honest failure modes. First, if your sales data is thin (under ten recorded calls, or a product with no sales conversation at all) your question harvest gets synthetic, and the models will happily invent plausible buyer questions that no buyer has ever asked. Substitute support tickets and forum threads, and accept that your stage classification will need a manual pass.

Second, teams publishing under about four pieces a month shouldn’t build 27 clusters. Build six, all in the two stages closest to purchase, and revisit in two quarters. A half-built decision map with the awareness half missing outperforms a complete map you’re three years from finishing.

Run the four-hour exercise on one product line this week. Compare the cluster list against whatever architecture you’re working from now, and count how many implementation-anxious clusters your current map has no home for. In every audit I’ve done, the answer is between six and twelve, and they are the pages your pipeline has been missing.