How Generative Engines Decide What to Cite
Most of what gets written about generative engine optimisation is ranking advice wearing a new hat. Build authority, earn links, match intent, be helpful. All fine, all several years old, and none of it explains why a 400-word page from a niche SaaS blog gets quoted in a Perplexity answer while your 3,000-word pillar, the one sitting at position two, gets nothing.
The reason is structural. Classic search returns a ranked list of URLs and lets the human do the reading. A generative engine has to do the reading itself, at speed, under a token budget, and then it has to attach a citation to a specific sentence it has just generated. Those are different jobs. Ranking answers “which pages are best for this query”. Citation answers “which chunk of text, from which corroborated source, can I lift and attribute without the answer falling apart”. Getting cited is a retrieval-and-selection problem. Once you see it that way, a lot of GEO advice stops making sense and a much shorter list of things starts to matter.
Your one keyword becomes fifteen queries before anything is retrieved
Ask ChatGPT Search a real question and expand the search step. You don’t see your target keyword. You see a fan-out: the model decomposes the prompt into sub-questions and issues several independent retrievals, often six to fifteen for a comparison or how-to prompt. Google’s AI Mode does the same thing, and their own patent filings describe it as query fan-out. Reconstructed from the visible search steps, a single prompt like “should a small content team use AI for briefs” expands to something close to this:
1. AI content brief tools comparison 2026
2. time saved using AI content briefs
3. AI brief quality problems small teams
4. content brief template AI prompt
5. do AI briefs hurt rankings
6. in-house content team size benchmarks UK
7. editor rejection rate AI drafts
Seven retrievals, seven candidate sets, and your page has to surface in at least one of them at the passage level. This is the first and most underrated implication for generative engine optimisation: you are not optimising a page for a query, you are supplying passages to an unpredictable set of sub-queries you never chose. A page that covers one topic thoroughly but only answers the head question in a self-contained way wins one lottery ticket. A page that answers seven adjacent sub-questions in seven clean, quotable passages holds seven.
The retrieval layer reads passages, not pages
Every one of these systems chunks documents before embedding them. Chunk sizes vary by vendor but the working range is roughly 200 to 800 tokens, split on headings and paragraph boundaries where the HTML makes that possible. Retrieval is hybrid: a lexical pass (BM25 or similar) plus a dense vector pass against an embedding model, then a reranker that scores each candidate passage against the sub-query. Cohere’s Rerank and similar cross-encoders are doing the final sort, and they score the chunk in isolation.
Here’s the consequence people miss. Your chunk arrives at the reranker with no title, no breadcrumb, no preceding paragraph, and no idea what “it” refers to. Think about what a mid-article paragraph looks like stripped of context:
CHUNK 47 (as scored, no surrounding context):
"As we covered above, this approach cut the time considerably.
The team found it worked best when paired with the framework
in the previous section, though results varied depending on
the factors mentioned earlier."
That chunk contains four anaphoric references and zero facts. No reranker will score it highly against any sub-query, because it is not about anything. Half the paragraphs in a typical well-written long-form post fail this test, because good long-form writing deliberately leans on what came before. Flow is a virtue for human readers and a tax in retrieval.
Corroboration is the filter between retrieval and citation
Retrieval gets you into the candidate pool. Selection decides who actually gets the footnote, and corroboration is the dominant signal there. Generative engines are built to reduce the risk of asserting something false with a citation attached, so the selection stage systematically favours claims that appear in multiple independently retrieved passages.
Watch this in practice. Run the same prompt through Perplexity five times and log the citations. You’ll typically see a stable core of three or four sources that appear in every run, plus a rotating tail. The stable core is almost always the set of sources whose claims agree with each other. Where a number is contested (say, two studies reporting different adoption rates), the engines either cite both and hedge the sentence, or drop the claim entirely and answer a safer adjacent question.
This inverts a piece of content marketing orthodoxy. Contrarian, entirely-original takes are excellent for building an audience and terrible for getting cited, because nothing corroborates them. The pages that get cited relentlessly tend to be ones that state a widely-supported fact precisely, with a number and a source, and then add their own angle afterwards. Put the corroborable claim in one passage and the hot take in the next one. The first gets retrieved and quoted; the second gets read by the human who clicks through.
Original research is the exception worth investing in, but only if it produces a claim other people will restate. A stat that gets picked up by three trade publications becomes corroborated within a quarter and then gets cited everywhere. A stat nobody repeats stays a single-source claim forever.
Extractability: can a sentence survive being lifted out?
The third factor is mechanical. The model needs a span it can quote or closely paraphrase, attached to an attributable claim. Run every important paragraph through four checks:
| Check | Fails when | Fix |
|---|---|---|
| Named subject | Passage opens with “it”, “this”, “the approach” | Repeat the entity name in every passage, even when it reads slightly repetitive |
| Claim in the sentence | The number lives in a chart, an image, or three sentences away | Put the figure in prose: “cut from 4.5 hours to 1.7 hours” |
| Self-contained scope | Qualifiers sit in a previous paragraph | Restate the scope inline: “for UK in-house teams of one to five” |
| Renderable without JS | Content injected client-side, or behind a tab or accordion | Server-render it. Fetch your URL with curl and check the claim is in the raw HTML |
That last one catches more sites than you’d expect. Tabbed comparison blocks, accordion FAQs, and anything loaded by a framework after hydration may be invisible to the crawlers feeding these indexes. Screaming Frog’s JavaScript rendering mode, run with rendering off, gives you the pessimistic view in about ten minutes for a 500-page site.
The rewrite, concretely
Take a real paragraph of the kind most of us ship:
We’ve seen this play out with clients. Their time to first draft dropped significantly after adopting the workflow described above, and quality held up too.
Nothing there is retrievable. No entity, no number, no scope, two backward references. Now the same claim, written for a reranker that will see it alone:
Small content teams using a structured AI brief cut time-to-first-draft from 4.5 hours to 1.7 hours across 38 briefs run between January and June 2026. Editor rejection rate stayed flat at 12%, which suggests the time saving did not come out of quality.
Same information, roughly the same length, and it now answers sub-queries 2 and 7 from that fan-out above. It also survives being quoted with a link next to it, which is the entire transaction you’re trying to complete. If you want the wider strategic frame around this, including how to prioritise which pages get the treatment first, our pillar on getting cited by AI search covers the programme-level view.
Measuring the thing you’re actually changing
Rank tracking will not tell you whether any of this worked. Two data sources will.
First, your server logs. The crawlers that feed these systems identify themselves, and the split between index-building bots and live-fetch bots tells you different things:
OAI-SearchBot/1.0; +https://openai.com/searchbot # index building
ChatGPT-User/1.0; +https://openai.com/bot # live fetch, user in session
PerplexityBot/1.0 (+https://perplexity.ai/perplexitybot)
Perplexity-User/1.0
ClaudeBot/1.0 (+claudebot@anthropic.com)
GPTBot/1.2 (+https://openai.com/gptbot) # training, not retrieval
A rise in ChatGPT-User and Perplexity-User hits on a URL means it’s being pulled into live answers, which is the leading indicator you want. GPTBot traffic tells you almost nothing about citation.
Second, a citation tracker. Profound, Peec AI, Otterly.ai and Semrush’s AI toolkit all sample prompts on a schedule and record which domains appear. Set up 30 to 50 prompts that mirror how your buyers actually ask, run them weekly, and track citation share rather than presence. Presence is noisy: these answers are non-deterministic, and a source that appears in one run of five is not really cited. Something appearing in four of five runs, consistently, over three weeks, is.
Pick your ten highest-intent pages this week and read them the way a reranker does: open each one, take any paragraph from the middle, paste it into a blank document with no title, and ask whether a stranger could tell you what it claims and who it’s about. The ones that fail are the ones quietly losing you citations you never knew were on offer.