What actually determines whether an AI assistant cites your brand
The observable signals behind citation selection, and which of them you can influence.
There is a specific kind of meeting we have been having a lot recently. A marketing director opens a laptop, types a question about their category into an AI assistant, and watches it recommend three competitors. Their own brand — larger, better known, with more traffic and a bigger content team — does not appear. The question that follows is always the same, and it is a fair one: why them and not us?
The honest answer is that nobody outside the model providers knows the full mechanism, and anybody who tells you otherwise is selling something. But that is not the same as saying nothing is knowable. Citation behaviour is observable. You can run the same query a hundred times, vary the phrasing, compare assistants, and watch which sources keep appearing. Do that for long enough across enough categories and patterns emerge — not a formula, but a set of tendencies consistent enough to act on.
This piece sets out what we have observed, what we think it means, and — more usefully — which of it you can do something about.
Retrieval is the part that matters
The first thing to understand is that when an assistant cites a source, two separate things have happened, and only one of them is really about your brand.
The first is retrieval. Faced with a question it does not want to answer purely from training data — anything current, anything commercial, anything where being wrong is expensive — the assistant issues one or more searches. Those searches go to a conventional index. The results that come back form a candidate set, usually a small one: a handful of pages, sometimes a dozen.
The second is selection. The model reads those candidates and decides which ones to lean on, summarise, and attribute. This is where the model's own judgement about clarity, authority and usefulness comes in.
If you are not in the candidate set, nothing else you do matters. Selection cannot rescue a page that retrieval never surfaced.
This is the single most consequential thing we can tell you, and it is also the least exciting. The dominant constraint on AI visibility is not some novel discipline. It is whether the search layer feeding the assistant can find you for the phrasings people actually use when talking to a machine.
Which is why a great deal of what passes for AI-era strategy collapses, on inspection, into search fundamentals that were true in 2015 and are still true now.
The questions people ask machines are different
That said, one thing genuinely has changed: the shape of the query. People type differently into a search box than they speak to an assistant. Search queries are compressed and keyword-like because two decades of using search engines taught everyone to strip out the grammar. Assistant queries are longer, more conversational, and far more likely to carry constraints and context.
Nobody types "best CRM" into an assistant. They type something closer to: we are a forty-person B2B services firm, our sales team hates the CRM we have, we need something that will not require a consultant to configure, what should we look at?
That query contains a company size, an industry, a failure state and a constraint. The retrieval layer will decompose it into several searches. The pages that surface are the ones that address those specific dimensions — not the ones optimised for the two-word head term.
The practical consequence: content built around constrained, situational questions gets retrieved for constrained, situational prompts. Content built around head terms competes in the most crowded part of an index for queries that assistants increasingly do not issue.
Signals we can see working
Across the categories we have looked at, the sources that keep getting cited tend to share a set of properties. None of these is a guarantee. All of them are within your control.
- They are actually crawlable. Content rendered entirely client-side, gated behind a form, or buried in a PDF is systematically under-represented. The retrieval layer is a crawler with an ordinary crawler's limitations.
- They answer the question near the top. Pages that state the answer in the first paragraph get cited more than pages that arrive at it after eight hundred words of scene-setting. A model extracting an answer will take the clearest available extract.
- They are structurally legible. Descriptive headings, short paragraphs, real lists, real tables. Not because a model loves markup, but because clean structure makes a confident extraction easy and ambiguous prose makes it risky.
- They are specific and falsifiable. Numbers, dates, named conditions, stated limitations. Vague marketing copy is difficult to summarise without the summary sounding like an advertisement, and assistants visibly avoid sounding like advertisements.
- They are corroborated elsewhere. A claim that appears only on your own site is a claim with one source. The same claim reflected in trade press, directories, review platforms and independent write-ups is a claim the model can restate with less risk.
- The entity is consistent. Same company name, same description, same category language across your site, your structured data, your profiles and your press coverage. Inconsistency splits the entity and dilutes everything attached to it.
Read that list again and notice something uncomfortable: there is nothing on it that a competent technical SEO would not have recommended a decade ago. That is the argument of this whole essay in one observation.
Third-party corroboration is doing more work than people expect
If we had to single out the most under-appreciated factor, it would be this one.
When an assistant answers a comparison or recommendation question, it very often does not cite vendor sites at all. It cites the reviews, the roundups, the trade coverage, the forum thread where six practitioners argued about it. From the model's perspective this is entirely rational: an independent source carries less risk of being an advertisement.
A great deal of AI visibility work is not work on your own website. It is work on what the rest of the internet says about you.
This is the point at which some clients become disappointed, because it is slower and less controllable than publishing a page. It involves getting into the roundups you are missing from, correcting the profiles that describe you wrongly, and giving trade press something worth writing about. It is closer to public relations than to content marketing.
It is also, in our experience, where the largest gaps usually are. Brands audit their own site obsessively and have never once checked what the eight independent pages ranking for their category question actually say about them.
What you cannot control
An honest account has to include the parts that are outside your reach.
Model providers strike commercial and licensing arrangements with publishers, and those arrangements shape which sources are readily available. Each assistant has its own retrieval stack and its own defaults, so the same question produces materially different citation sets across products. Model updates shift behaviour without notice; a set of queries you were cited on in one month can look different the next, with no change on your side.
There is also an unavoidable amount of variance. Ask the same question twice and you can get two different citation sets. Any measurement approach that treats a single check as a result is measuring noise. Sampling repeatedly across time is the only way to see a real signal.
The correct response to all of this is not despair; it is to stop treating citation as a rank you can hold and start treating it as a probability you can shift.
How to actually work on this
The sequence we use, in the order we use it:
- Establish a baseline. Build a set of the questions a real buyer would ask an assistant in your category — not keywords, questions. Run them repeatedly, across assistants, over time. Record who gets cited and what is said about you when you are mentioned.
- Fix retrieval before anything else. Crawlability, rendering, indexation, structured data, internal linking, page speed. Unglamorous, and it determines the ceiling for everything after it.
- Audit the third-party layer. For every question where you are absent, look at what is being cited instead. Those pages are your actual target: get into them, get corrected in them, or earn something better.
- Write for the constrained question. Publish pages that address the specific situations buyers describe, and answer them in the first paragraph.
- Re-measure on a schedule, not on a hunch. Monthly is usually enough. Anything more frequent and you are reading variance as progress.
The uncomfortable conclusion
The reason so many companies are spending on AI without seeing anything back is that they bought capability and left distribution alone. They deployed assistants internally, ran pilots, licensed tools — and made no change whatsoever to whether an assistant recommending options in their category would ever mention them.
Getting cited is not a new discipline requiring new vendors. It is the old discipline, applied to a surface where the results page happens to be a paragraph of prose instead of ten blue links. The brands doing well at it are, almost without exception, the ones that were already doing the boring parts properly.
If you want to know how your brand currently appears across assistants for the questions your buyers are actually asking, that is what our AI Visibility Audit measures.
