What 200,000 AI Citations Reveal About Which Sites Get Cited

Large-scale citation research shows AI answers concentrate on a narrow set of domains, favour question-led formats and reward established pages.

AB
Aanchal BhatiaSEO Strategist
Explore this article in ChatGPTExplore this article in ClaudeExplore this article in Perplexity
Citation tally sheet: four source types carry dozens of marks each, while a long tail of domains carries one mark.

Key Highlights

  • Citation across AI answer engines follows a steep concentration curve: a narrow set of domains is cited repeatedly, while an enormous tail of domains is cited once or twice and never again.
  • Independent auditing has found that a meaningful share of the sources these engines cite show evidence of being AI-generated themselves, which changes what "being cited" actually proves.
  • Question-led formats, direct comparisons and structured how-to content are extracted far more readily than narrative essays, because they contain passages an engine can lift whole.
  • Cited pages skew established rather than freshly published, which makes maintenance of existing assets a higher-return activity than constant new publishing.
  • Breaking into the repeatedly cited group is realistic within a narrow topic and close to impossible across a broad one, so scope discipline beats volume.
  • Citation share, not raw traffic, is the metric that tells you whether any of this is working, and it needs a baseline before you start changing pages.

Most advice about getting cited by AI engines is somebody's hunch dressed up as a principle. It sounds plausible, it spreads quickly, and almost none of it has been checked against what these systems actually do at scale. Large-scale citation analysis is the correction. When you stop looking at one answer and start looking at hundreds of thousands of them, the individual noise cancels out and the underlying structure becomes visible, and that structure frequently contradicts the confident guidance circulating in your feed.

According to an audit of four generative search engines published on arXiv by researchers at Northwestern University, an analysis of 712 real-world queries across politics, health and the environment found evidence of AI-generated sources being cited across all four engines, at roughly 16% of cited sources. That single finding reframes the whole exercise. The citation pool is not a clean ranking of the web's most trustworthy pages. It is a mixture, and part of what these systems cite is machine-written content that happened to be positioned well.

This piece works through what the citation research genuinely establishes and what it does not. It covers why aggregate studies beat anecdotes, how brutally concentrated citation actually is, which formats get extracted most often, how old the typical cited page is, and how to turn all of that into a plan you can act on this quarter rather than a set of observations you nod along to and forget. Our own smaller-scale teardown of 500 AI search citations is a useful close-up companion to the broad patterns described here.

Why Do Large-Scale Citation Studies Matter More Than GEO Opinion?

Large-scale studies matter because they show aggregate behaviour across thousands of queries, which no single observation can. One screenshot of an AI answer tells you what happened once. A study across hundreds of thousands of citations tells you what tends to happen, and only the second thing can support a strategy.

The problem with anecdote-driven optimisation advice is not that it is dishonest. It is that a single AI answer is an extremely unreliable sample. These systems show meaningful run-to-run variation, so the same question asked twice can return different sources. Anyone drawing a rule from one observed answer is reading noise as signal, and then publishing the noise as a tactic that other people adopt.

Aggregate research fixes the sample problem. When a study examines citations across many queries, many topics and several engines, individual variability averages out and genuine tendencies surface. You get to see, for instance, that citation concentrates rather than spreads, which is a structural fact about the system rather than an accident of one query.

What these studies give you is directional confidence, not a formula. They tell you which broad bets are supported by observed behaviour and which are not. You still have to adapt the direction to your own niche, your own audience and your own content, because no published study was run on your specific topic. The honest framing is that the data narrows the search space considerably, and then your own measurement narrows it the rest of the way.

There is a second, less obvious benefit. Knowing the aggregate pattern protects you from expensive mistakes. If the research shows that citation rewards established, maintained pages, then a plan built on publishing thirty new posts a quarter is fighting the observed behaviour of the system. Recognising that before you commit a year of budget is worth considerably more than any individual tactic the same research might suggest. It is also why the underlying question of whether AI visibility is more about trust than rankings repays study more than any single tactic does.

How Concentrated Is AI Citation Across Domains?

Infographic showcasing the steep concentration curve of AI citation, with a narrow repeatedly cited core of domains against an enormous tail of domains cited only once or twice
Two groups, not one distribution — and only one of them is worth entering.

Citation is heavily concentrated. Across audited engines, a comparatively narrow group of domains is cited again and again, while a very large number of domains appear only once or twice. The practical effect is that most sites are technically citable but functionally invisible, because a single citation in a niche query changes nothing.

The Northwestern audit is unusually direct about this shape, and it is worth reading the finding in the researchers' own words rather than through someone's summary of it.

"We observed that generative search engines include a somewhat narrow set of repeatedly cited domains while predominantly surfacing a large number of minimally cited domains in responses to users' queries." Mowafak Allaham and Nicholas Diakopoulos, authors, Synthetic Sources? Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources. Source: arXiv

Two separate groups are described there, and confusing them is where most strategy goes wrong. There is a repeatedly cited core, which is small and hard to enter. Then there is a long tail of minimally cited domains, which is enormous and easy to enter. Appearing once in that tail feels like progress and delivers almost nothing, because the query that surfaced you may never be asked in that form again.

This is the uncomfortable part of the research for smaller publishers, and it deserves stating plainly rather than softening. Being good is not sufficient. The repeatedly cited core is populated by domains that engines have reason to return to across many different questions, which usually means recognised reference sources, established publications and organisations with obvious subject authority. Competing with that across a broad topic is not a realistic plan for most sites.

What is realistic is narrowing the field until the concentration works for you instead of against you. Within a genuinely specific topic, the pool of plausible sources is much smaller, and the repeatedly cited core for that topic may contain only a handful of domains. That is a group you can join. The strategic instruction the data supports is not "make better content" but "make the topic smaller until repeated citation is achievable, then own it".

It is also worth being honest about what concentration does not mean. It does not mean citation is fixed or that incumbents are permanent. The audited engines vary in which domains they favour, and coverage shifts as content changes. Concentration describes the shape of the distribution, not a locked ordering, and shapes can be entered at the right scale.

What Does the Citation Pool Actually Contain?

Infographic showcasing what an AI citation actually proves, with roughly one in six cited sources showing evidence of being AI-generated and the levers of citation reframed as mechanical rather than reputational
A citation proves retrievability and match — not that anyone vouched for you.

The pool is not a curated list of authoritative pages. The four-engine audit found roughly 16% of cited sources showed evidence of being AI-generated, meaning a noticeable fraction of what these systems present as sourcing is itself machine-written content that was well positioned rather than independently authoritative.

This finding deserves more attention than it usually gets, because it changes what a citation proves. If you assume the citation pool is a quality ranking, then earning a citation feels like external validation of your expertise. If a meaningful share of cited sources are themselves synthetic, then a citation is better understood as evidence that your page was retrievable, extractable and topically well matched at the moment the question was asked.

That reframing is genuinely useful rather than merely deflating. It tells you the levers are mechanical as much as reputational. Retrievability, clear structure, unambiguous topical match and the presence of directly quotable passages are things you control. They are not proxies for how respected you are, which is precisely why a smaller site can compete on them, and why AI engines regularly cite pages that do not rank on Google at all.

It also introduces a competitive risk worth naming. If synthetic content can enter the citation pool, then in any topic where mass-produced content is cheap to generate, the tail gets noisier and the value of being merely present drops further. The defensible position is the one that thin content cannot imitate: original data, direct experience, specific numbers, named methods and details that come from having actually done the thing.

There is a governance dimension here too, which the researchers frame as a risk to users rather than an opportunity for marketers. Their concern is that people treat synthesised answers as equivalent to authoritative sourcing. For anyone producing content, the reasonable response is to be the kind of source that holds up when this gets scrutinised more carefully, because the direction of travel is clearly towards more scrutiny of source quality, not less.

Which Content Formats Get Cited Most Often?

Infographic showcasing why question-led, comparison and step-by-step formats get extracted most often by AI engines, contrasting a self-contained passage with one that depends on surrounding build-up
Format advantage is really extraction advantage.

Formats that answer a specific question directly are cited most: question-and-answer content, side-by-side comparisons, and structured step-by-step instructions. These formats contain self-contained passages an engine can extract without reconstructing an argument, which makes them cheap to cite and easy to match to a query.

The mechanism matters more than the list. An answer engine is assembling a response from fragments, so what it needs is a passage that stands on its own. A paragraph that only makes sense after three paragraphs of build-up is expensive to use. A paragraph that opens by directly answering the question it sits under is cheap to use. Format advantage is really extraction advantage.

Comparison content performs particularly well because comparative questions are extremely common and genuinely hard to answer from a single source. When someone asks which of two options suits a particular situation, a page that lays out the trade-off explicitly, with the criteria stated, is doing work the engine would otherwise have to do by stitching together several sources. Explicit trade-offs, stated conditions and clear recommendations are what get lifted.

Instructional content earns citations for a similar reason. Sequential steps have unambiguous boundaries, so an engine can take steps three to five without mangling the meaning. Vague process writing that buries the actual instruction inside narrative does not offer those boundaries, which is why two pages covering the same procedure can perform completely differently.

The practical translation is structural, and it is mostly about the first sentence under each heading. Phrase headings the way the question is actually asked. Answer immediately underneath in a passage that would still make sense if someone pasted it somewhere else on its own. Keep one idea per section so the extractable unit is clean. Then let the supporting detail follow, for the human reader who wants the reasoning rather than the answer.

Also read: How to optimise content for AI search, which covers the structural changes that make a page extractable rather than merely thorough.

How Old Is the Typical Page That Gets Cited?

Cited pages skew established rather than newly published, frequently a year old or more. Engines favour content that has had time to accumulate references and corroboration, which means citation rewards durability and maintenance far more reliably than it rewards being first to publish on a topic.

This cuts directly against a publishing instinct that most content teams have internalised. The assumption is that fresher is better and that a steady stream of new posts keeps you visible. In answer engines the assumption largely fails, because these systems are not trying to surface the newest take on a subject. They are trying to give a settled, defensible answer, and settled answers tend to live on pages that have been around long enough to be corroborated elsewhere.

There is an important exception worth stating so the principle is not over-applied. For genuinely time-sensitive questions, where the correct answer changed recently, recency matters enormously and old pages actively hurt. The rule is not "old content wins". It is that stability wins for stable questions, and most commercially valuable questions are stable ones.

The strategic consequence is a reallocation of effort rather than a reduction of it. If durable pages earn citations, then the highest-return work is usually improving what you already have: deepening a page that half-answers a question, consolidating several thin overlapping posts into one substantial resource, updating figures that have gone stale, and adding the specific detail a thin version skipped. That work compounds on an asset that is already accruing trust, whereas a new post starts from zero.

It also changes how you should judge a page's performance. A page published six weeks ago that has not been cited is not a failed page, it is an immature one. Judging AI visibility on the same timescale you would judge a social post produces exactly the wrong decisions, usually deleting or rewriting pages that simply had not aged into the citation pool yet.

Why Do Reference and Community Sources Appear So Often?

They appear often because they contain specific, first-hand, question-shaped material at enormous scale. A forum thread answering a narrow practical question, or a reference entry stating a definition plainly, is exactly the sort of self-contained, directly responsive passage these systems extract most readily.

The lesson usually drawn from this is that you should go and post on those platforms. That is only half right, and the half that is right is easily overdone. Participating in the places where your audience genuinely discusses your subject is legitimate and can put useful material into circulation. Doing it transparently, as a knowledgeable participant rather than a promoter, is the only version that survives contact with those communities.

The more durable lesson is about what makes that material citable in the first place. These sources win on specificity. They answer the actual narrow question someone asked, with concrete details, rather than the broad topic that question belongs to. Most brand content does the opposite: it covers a topic comprehensively and answers no single question crisply. You can adopt the specificity without adopting the platform.

There is a defensive angle too. If third-party discussion is influencing what engines say about your category, then knowing what that discussion contains matters, because it may be shaping answers about you that you never see. That is an argument for monitoring how you are represented, not just whether you appear. It also explains a pattern many teams find baffling, where a page ranks on page one of Google yet never appears in an AI overview because it answers the topic rather than the question.

How Do I Measure My Own Citation Position?

You measure it by tracking citation share on a fixed set of questions across engines: how often you appear, on which questions, and against which competing domains. Traffic figures cannot tell you this, because an answer that cites you may generate no visit at all while still shaping the reader's view.

Start by fixing the question set, because everything else depends on it. Build a list of the questions that genuinely matter commercially, phrased the way a person would actually type them rather than as keywords. Keep the list stable over time. A moving question set produces numbers that cannot be compared month to month, which is the most common way this measurement quietly becomes useless.

Then record, for each question and each engine, whether you are cited at all, whether you are cited among the first sources or buried further down, and which other domains appear alongside you. That competitive column is the one people skip and the one that carries the most information, because it shows you who occupies the repeatedly cited core for your specific topic, and therefore who you actually have to displace.

Check the same questions across the major engines rather than one, since they differ in what they favour. Running the set through ChatGPT, Perplexity, Google Gemini and Microsoft Copilot will usually show noticeably different source patterns for identical questions, and a strategy tuned to one engine can be invisible in another.

Supplement this with whatever first-party reporting exists. Bing Webmaster Tools has begun surfacing AI-related performance reporting for sites, which gives you a platform-side view rather than an observational one. Treat it as a useful cross-check on your manual tracking rather than a replacement for it, since coverage is still partial.

Two habits keep this honest. Re-run the whole set on a schedule rather than checking when you feel curious, because these systems vary run to run and one-off checks mislead in both directions. And record what you changed and when, so that a movement in citation share can be attributed to something rather than admired as a mystery.

Also read: Six metrics that actually matter for measuring AI search visibility, which sets out how to define each one before you start tracking.

What Should I Change First Based on This Research?

Infographic showcasing the ordered response to the citation research — narrow the scope, restructure existing pages, then maintain them as compounding assets — with the citation-share measurement loop that verifies each change
Scope, then structure, then maintenance — measured by citation share, never traffic.

Change scope first, then structure, then maintenance. Narrow the topic until repeated citation is plausible, restructure your strongest existing pages so each section answers one question in a self-contained passage, and put your best pages on a maintenance schedule instead of publishing new ones.

Scope comes first because it determines whether the rest of the work can succeed at all. Write down the specific area in which you could plausibly become one of a small number of repeatedly cited sources. If that description could equally describe fifty other companies, it is still too broad. Narrow it by audience, by use case, by industry or by the specific decision you help people make, until the honest answer is that a short list of sources covers it and you could join that list.

Structure comes second because it is the cheapest meaningful improvement available. Take the pages that already cover commercially important questions and rework the headings into the questions people actually ask, then make the first passage under each heading answer it completely. This is usually a few hours per page and it changes how extractable the page is without changing what it says.

Maintenance comes third and matters most over time, because the research points to established pages as the ones that get cited. Decide which pages are your genuine assets, then commit to revisiting them on a schedule: refreshing figures, adding detail that was missing, absorbing thinner posts that overlap. A maintained page keeps compounding, and the alternative, an ever-growing library of pages nobody revisits, dilutes attention across assets that never mature.

Underneath all three sits the thing that cannot be shortcut, which is having something specific to say. Original data, real methodology, actual numbers from your own work, named constraints and honest trade-offs are what make a page worth returning to when synthetic content is abundant and cheap. If you are tracking where you stand across engines and want a starting baseline, Rank in AI Overview runs a free AI-visibility check that maps which questions currently cite you and which cite someone else.

What Do People Get Wrong When Reading Citation Studies?

The most common errors are treating aggregate findings as guarantees for a specific niche, copying the platforms that get cited rather than the properties that make them citable, and measuring success by appearing once rather than by appearing repeatedly on questions that matter commercially.

Over-generalising is the first trap. A study run across politics, health and environmental queries describes those domains. The structural findings, such as concentration and format preference, probably travel well because they follow from how the systems work. The specific proportions may not travel at all. Treat the shape of the finding as portable and the exact numbers as local to the study.

Platform mimicry is the second. Noticing that community and reference sources get cited and concluding that you should imitate their format wholesale misreads the cause. Those sources win on specificity and question-shaped answers, not on being forums or encyclopaedias. Copy the property, not the platform.

Vanity counting is the third and probably the most costly. A dashboard showing that you were cited somewhere last month is not a measurement, because a single citation on an obscure question has no commercial consequence. What matters is repeated citation on the questions that precede a purchase decision, which is a much smaller and much harder number to move.

The fourth error is impatience. Because cited pages skew established, most of the work described here has a lag between the change and the visible result. Teams that re-evaluate after three weeks routinely abandon correct strategies. Set the review window to match the timescale on which citation actually shifts, which is months rather than weeks.

Conclusion

Large-scale citation research does not hand you a formula, and anyone selling one from it is overreading the data. What it does give you is an accurate picture of the terrain: citation concentrates on a narrow repeatedly cited core, an enormous tail of domains appears once and disappears, formats that answer questions directly get extracted most, established pages outperform new ones, and a measurable share of what gets cited is machine-written rather than authoritative.

Each of those findings points the same way. Narrow your scope until repeated citation is achievable, structure pages so each section answers one question in a passage that stands alone, maintain your genuine assets instead of accumulating new ones, and judge progress by repeated citation on commercially meaningful questions rather than by any single appearance. None of that is fast, which is exactly why most competitors will not do it.

Rank in AI Overview researches how answer engines select and cite the sources they trust, and is building an AI visibility tool arriving shortly. If you want to know where you currently sit in that distribution, a free AI visibility audit will show which questions cite you today and which cite your competitors instead.

Frequently asked questions

Which websites get cited most often by AI engines?+

Audits consistently find a narrow group of domains cited repeatedly, alongside a very large tail of domains cited only once or twice. Recognised reference sources, established publications and organisations with clear subject authority dominate the repeatedly cited group.

Does a citation prove my content is authoritative?+

Not reliably. Research auditing four engines found roughly 16% of cited sources showed evidence of being AI-generated, so citation indicates that a page was retrievable, well structured and topically matched, rather than independently verified as authoritative.

What content format earns the most citations?+

Question-and-answer content, direct comparisons and step-by-step instructions perform best. They contain self-contained passages an engine can extract without reconstructing surrounding context, which makes them substantially cheaper to cite than narrative writing.

Should I publish more content or improve existing pages?+

Improve existing pages in most cases. Cited pages skew established rather than new, so deepening, consolidating and maintaining your strongest assets generally returns more than adding fresh posts that start with no accumulated corroboration.

Can a small website realistically get cited?+

Yes, but only within a narrow topic. Broad categories are dominated by a small repeatedly cited core that is very hard to enter. Within a specific subject the plausible source pool is small enough that focused authority can win.

How long before changes affect my citation position?+

Expect months rather than weeks. Because engines favour established, corroborated pages, structural and depth improvements need time to be recrawled, referenced and settled before citation behaviour reflects them. Reviewing after a few weeks produces misleading conclusions.

Want more of RankAI?

One playbook a week. Tactical, no fluff.

Join the waitlist
Continue reading

Related articles