AI Search Visibility

How AI Search Engines Work: Retrieval, Fan-Out, and Citations

Direct Answer

AI search engines do not simply look up one keyword and return one ranked list. They interpret the user's prompt and context, decide whether current information is needed, often generate several related searches, retrieve candidate pages or excerpts, select evidence, synthesize an answer, and decide which sources to show. That last step is probabilistic: the same prompt can produce a different answer or citation set on another run.

Sam Oh's first full lesson in the Ahrefs AEO course gives marketers a useful mental model: AI answers can draw on learned model knowledge and information retrieved in real time. Search optimization still matters because retrieval systems need accessible, relevant, authoritative pages. The new work is understanding one-to-many query fan-out, making evidence easy to extract, and measuring visibility as a distribution rather than a fixed rank.

The durable model: prompt and context → search decision → query fan-out → retrieval → evidence selection → synthesis → visible citations. Different products can skip, combine, or repeat these stages.

Watch the Ahrefs Lesson

Video and course credit: Sam Oh and Ahrefs. Oh presents this 8-minute lesson as part of Ahrefs' free AEO course. This article independently organizes the mechanics, updates statistics that changed after the recording, and adds a repeatable visibility-testing workflow. Ahrefs' studies are observational and product-specific, not disclosures of universal ranking factors.

The AI Search Pipeline, Step by Step

No outsider can see the complete internal pipeline for every answer engine, and the products change quickly. Still, the following model is useful because it separates stages that marketers often collapse into a single word: "ranking."

StageWhat the system may doWhat publishers can influence
1. InterpretResolve intent, entities, constraints, conversation history, location, date, and requested format.Use unambiguous names, facts, scope, and language that matches real customer questions.
2. DecideAnswer from available context or invoke search, files, connected apps, databases, or other tools.Keep current information publicly accessible where appropriate, with stable URLs and clear update dates.
3. Fan outTurn one prompt into several related searches, sometimes in parallel.Cover the decision journey and connected subtopics, not only one head term.
4. RetrievePull results, snippets, passages, structured facts, or document chunks from one or more indexes.Make each important section understandable when read apart from the whole page.
5. SelectFilter, rerank, diversify, and choose evidence that appears useful for the answer.Publish specific claims, primary evidence, explicit entities, and clear limitations.
6. SynthesizeCombine evidence with model reasoning into a response appropriate to the user's request.Keep facts consistent across owned pages and earn independent corroboration.
7. CiteAttach some retrieved sources to claims or show a source list, depending on product policy.Make the page a strong, directly relevant support for the exact claim it should substantiate.
8. ContinueUse follow-up questions, memory, or additional searches to refine the answer.Anticipate comparisons, objections, implementation questions, and evidence requests.

The interface can hide most of this work. A user sees one answer, while the system may have searched several variations, inspected partial excerpts, discarded candidate sources, and run another retrieval pass. That is why a traditional rank tracker cannot fully describe AI visibility.

Training Data and Retrieval Are Different Sources of Context

Sam divides an AI answer's information into two broad sources. The first is parametric knowledge: patterns and information learned during model training. The second is retrieved context: current pages, files, app data, or tool results supplied while the answer is being generated.

Context sourceStrengthImportant limit
Training dataBroad learned knowledge and language capability without a live search.It is not a guaranteed factual database, can be outdated, and may not preserve source provenance.
Web retrievalCurrent or niche information from searchable public sources.Retrieval quality depends on the index, query, access, excerpt, reranking, and product policy.
Conversation and memoryUser-specific goals, prior turns, preferences, and constraints.It can personalize results, which makes two users' answers less comparable.
Files and connected appsPrivate, first-party context such as documents, email, analytics, or a knowledge base.Permissions, freshness, and data quality matter; this evidence may never be public or citable.
Tools and databasesStructured calculations, product data, maps, finance, weather, code, or specialist APIs.Tool availability and source disclosure vary by system and plan.

The transcript suggests that training snapshots may update about every six months. That is a useful illustration, not a universal cadence. Providers train, fine-tune, refresh, and ground models on different schedules. For AEO work, the reliable question is not "When was the model trained?" It is "What current evidence can this product retrieve for this prompt, under these conditions?"

What RAG Means, and What It Does Not

Retrieval-augmented generation, or RAG, was formalized as a way to combine a model's parametric memory with non-parametric retrieved memory. In plain English, the system fetches relevant material and places it in the model's working context before the model writes the answer.

Modern AI search is broader than a textbook RAG diagram. A commercial product may use its own search index, a search partner, multiple retrieval tools, ranking models, document chunking, citation alignment, conversation memory, and safety filters. Saying that every answer engine simply "calls a web search API" is therefore too narrow.

RAG does not guarantee truth. The system can retrieve the wrong passage, miss a critical source, misunderstand context, combine incompatible claims, or write a sentence that exceeds what the cited page supports. Retrieval improves grounding; it does not eliminate verification.

Query Fan-Out Turns One Prompt Into Many Searches

A conventional keyword search is often modeled as one query producing one result page. AI search can be one-to-many. A request such as "plan a first trip to Japan in October for two people who dislike crowds" may trigger searches about weather, neighborhoods, rail passes, regional events, quieter destinations, opening hours, and itinerary sequencing. The system then combines those findings into one response.

The numbers deserve precision. A Nectiv analysis of more than 60,000 Google fan-out queries found that 59% of observed results contained between five and 11 fan-outs, with a maximum of 28. That does not mean nine to 11 is a universal average. Ahrefs also documented a single ChatGPT Deep Research demonstration in which a red phone-case request produced 420 searches and 30 cited sources. It is an extreme example of a research workflow, not a normal count for every AI answer.

Why fan-outs are not a new keyword list

Ahrefs reports that more than 95% of observed fan-out queries had no recurring search volume in its database. They are synthetic, context-rich, and can change between runs. Their value is diagnostic: repeated subtopics reveal what evidence the answer engine may need to complete a decision.

  • Cluster fan-outs by need: definition, comparison, fit, price, risk, proof, implementation, and alternatives.
  • Look for recurring evidence gaps across several prompts and engines.
  • Improve or consolidate the best existing page before creating another URL.
  • Keep one topic coherent enough that passages still make sense when retrieved separately.
  • Do not generate a thin page for every synthetic variation.

Google's official guidance for AI search experiences makes the last point explicit: foundational SEO practices still apply, and scaled pages created mainly to manipulate rankings or AI responses can violate spam policy.

Retrieval, Use, Citation, Mention, and Click Are Five Different Events

A page does not move through a guaranteed funnel. Each stage can stop independently.

EventWhat it meansWhat it does not prove
RetrievedThe system fetched the page, snippet, or passage as a candidate.That the answer used it.
Used as evidenceInformation from the source influenced the generated answer.That the source will be visible to the user.
CitedThe interface linked or attributed a claim to the source.That the user noticed or clicked it.
MentionedThe answer named a brand, product, or person.That an owned page was retrieved or cited.
ClickedThe user followed a link or source card.That the visit was qualified or converted.

This is why "AI rank" is a weak standalone metric. Ahrefs' own technical guide describes these systems as probabilistic and recommends looking at average visibility across multiple prompts. A practical dashboard should separate mention frequency, citation frequency, cited URLs, factual accuracy, referral sessions, assisted conversions, and direct customer attribution.

The 76% Citation Statistic Is Already Outdated

The lesson cites Ahrefs' July 2025 finding that 76.1% of Google AI Overview citations came from pages ranking in the top 10 for the original query, 9.5% came from positions 11 to 100, and 14.4% did not rank in the top 100. It supported a sensible point: traditional search visibility and AI citations overlap.

Ahrefs repeated the analysis in March 2026 with 863,000 keyword result pages and about four million AI Overview URLs. The distribution changed dramatically. For standard blue-link results, 37.1% of citations ranked in the top 10, 26.2% ranked between 11 and 100, and 36.7% did not rank in the top 100 for the original query.

Ahrefs studyTop 10Positions 11-100Outside top 100Interpretation
July 2025
1.9M citations
76.1%9.5%14.4%Strong overlap with the original query's top results.
March 2026
standard blue links
37.1%26.2%36.7%Fan-out searches and broader retrieval appear to surface many sources outside the original SERP.

These studies used different snapshots and the search product changed between them, so do not splice them into a trend line or universal rule. The useful conclusion is narrower: ranking well can improve retrievability, but the original prompt's top 10 is neither necessary nor sufficient for citation. Topic coverage, passage relevance, fan-out visibility, freshness, diversity, and source selection all sit between ranking and a visible link.

Freshness, Consensus, and Authority Need Nuance

Freshness

Ahrefs analyzed almost 17 million cited URLs in 2025 and found that AI-cited pages were 25.7% newer by publication date than top organic results on average. By last-update date, the difference was 13.1%. The platform split matters: ChatGPT skewed toward newer pages, while Google's AI Overviews looked much closer to organic search.

Treat freshness as query-dependent evidence, not permission to change a date. News, prices, product specifications, software, laws, and recommendations need current support. Historical explanations may not. Update facts, examples, sources, and recommendations when reality changes, then make the revision visible.

Consensus

Consistent descriptions across independent sources can make an entity easier to represent. But "consensus" is not a publicly documented dial that marketers can turn. Repeating a claim across owned pages or manufacturing community mentions is not independent corroboration. Better inputs include primary research, customer evidence, expert commentary, public demonstrations, reputable directories, and editorial coverage that exists because it helps the audience.

Authority

Classic SEO still helps because crawlable, linked, well-ranked pages are easier for retrieval systems to find. Yet authority cannot rescue an irrelevant passage. AI systems may work from titles, snippets, and isolated chunks. A useful page makes the relevant entity, answer, qualification, evidence, and date clear inside the section that may be retrieved.

A Practical Topic-Coverage Workflow

1. Choose prompts tied to decisions

Select five questions that precede a sale, shortlist, risk review, or support decision. Avoid vague vanity prompts such as "What are the best companies?" Add audience, market, constraints, and intended outcome.

2. Run controlled repetitions

Test each prompt five times in two relevant systems. Record the date, model or product, location, login state, conversation context, and whether search was visibly triggered. Use clean sessions when measuring a baseline, because memory and prior turns can personalize answers.

3. Extract recurring evidence needs

Capture cited URLs, named competitors, factual claims, visible subqueries when available, and recurring subtopics. Cluster them by the buyer's decision journey. Do not assume a hidden fan-out exists merely because a related phrase appears in the answer.

4. Map evidence to existing pages

For each recurring need, identify the best owned page, the primary evidence it contains, relevant third-party support, and the gap. Fix contradictions and merge overlapping pages before adding more inventory.

5. Improve one authoritative source

Add a direct answer, decision criteria, dated evidence, limitations, original examples, author or organization context, and primary references. Write sections that remain clear when retrieved alone. Keep the page useful to a human who arrives from a citation.

6. Retest the distribution

Repeat the same protocol after the page is recrawled. Compare mention and citation frequency across runs, not one favorable screenshot. Then connect the visibility data to referrals, conversions, assisted attribution, and factual accuracy.

A good AEO deliverable is not "we ranked number one in ChatGPT." It is a documented prompt set, source map, evidence gap, page revision, repeated result distribution, and business outcome.

Copy-Ready AI Search Evidence Test

Objective: measure how answer engines represent and source one decision topic.

Business/entity: [canonical name and URL]
Audience and market: [specific buyer, country, language]
Decision: [purchase, shortlist, implementation, risk review]
Platforms: [two systems your audience uses]
Test date range: [dates]

Prompts:
1. [prompt with audience, constraints, and outcome]
2. [prompt]
3. [prompt]
4. [prompt]
5. [prompt]

For every prompt, run five clean-session tests per platform and record:
- platform, model/product, date, location, login state
- whether live search was triggered or visible
- recurring subtopics or visible fan-out queries
- brand mentioned: yes/no and surrounding claim
- cited domains and exact URLs
- whether each citation supports the linked claim
- material factual errors or missing qualifications
- competitor mentions and cited evidence
- owned page that should answer the gap

Report:
1. Mention frequency by prompt and platform.
2. Citation frequency and source diversity.
3. Unsupported or inaccurate claims.
4. Recurring evidence gaps grouped by decision stage.
5. Three page improvements ranked by buyer value and evidence strength.
6. Referral, conversion, and self-reported attribution where available.

Rules:
- Do not treat one response as a stable rank.
- Do not invent hidden fan-outs, search volume, or retrieval events.
- Distinguish retrieved, cited, mentioned, clicked, and converted.
- Label observations separately from causal claims.
- Recommend consolidation before creating a new page.

Source-Readiness Checklist

CheckPass conditionEvidence
Direct answerThe relevant section answers the question before expanding.First paragraph under the matching heading
Entity clarityNames, products, locations, dates, and relationships are explicit.Visible copy plus appropriate structured data
Passage independenceA retrieved paragraph or table remains understandable without the entire article.Section-level review
Primary evidenceImportant claims link to the original documentation, dataset, or first-party record.Source map and claim audit
Current factsVolatile claims show a checked date and genuine update history.Change log or editorial record
Useful distinctionThe page contributes data, experience, examples, or synthesis beyond paraphrasing other pages.Original artifact, method, or analysis
Technical accessThe canonical page returns 200, renders key content, is internally linked, and is not unintentionally blocked.Crawl, index, robots, and canonical checks
Human destinationA user who follows the citation can verify the claim and take a sensible next step.Page QA and conversion path

Video Chapters

TimeTopic
00:00Why AI search mechanics matter for AEO
00:30Training data and real-time retrieval
01:19RAG and grounding with web sources
02:00Two routes to influence AI answers
02:34How query fan-out changes search
03:28Fan-out counts and the Deep Research example
04:24Viewing fan-outs in Ahrefs Brand Radar
04:36Why synthetic fan-outs are not keywords
05:14How AI systems choose citations
05:26Probabilistic answers and visibility
06:13Consensus, freshness, and authority
06:45The original 76% top-10 study
07:03Citations from outside the top 100
07:24The complete AI search mental model
08:09Platform differences and the next lesson

Bottom Line

Sam Oh's lesson provides the right foundation: AI search combines learned model knowledge with retrieved evidence, and query fan-out expands one prompt into a wider research problem. Traditional SEO remains important because searchable, authoritative pages feed retrieval. But citation is a later, probabilistic selection event, not a renamed blue-link ranking.

The practical move is not to chase every synthetic fan-out. Build the smallest set of pages that fully supports a customer's decision, make each important section independently understandable, publish real evidence, keep volatile facts current, and test a fixed prompt set repeatedly. Measure how often the brand is accurately represented and supported, then connect that visibility to qualified behavior.

Sources and Link Map

Common questions

How do AI search engines work?
An AI search system interprets the prompt and context, decides whether fresh retrieval is needed, may generate multiple related searches, retrieves candidate documents or excerpts, selects evidence, synthesizes an answer, and may display citations. The exact pipeline differs by product, query, model, index, and user context.
What is retrieval-augmented generation?
Retrieval-augmented generation, or RAG, combines a generative model with information retrieved at answer time. The retrieved material becomes grounding context for the response. Modern commercial search products can add reranking, tools, proprietary indexes, conversation memory, and citation systems beyond the original academic RAG pattern.
What is query fan-out?
Query fan-out is the process of breaking one user prompt into several related searches that can run in parallel. The resulting evidence is merged into one answer. Fan-outs are often synthetic and variable, so they are better used to identify recurring topic gaps than as a fixed keyword list.
Should I create a separate page for every fan-out query?
No. Group recurring subquestions by user need and improve the smallest set of authoritative pages that covers the decision journey. Google explicitly warns that creating scaled pages mainly to manipulate rankings or AI responses can violate its spam policy.
Is an AI citation the same as a search ranking?
No. Generated answers and source choices are probabilistic. A page can be retrieved without being visibly cited, cited without receiving a click, or mentioned without a linked citation. Measure frequency across repeated tests rather than treating one response as a permanent rank.
Does a Google top-10 ranking guarantee an AI Overview citation?
No. Ahrefs found meaningful overlap, but its March 2026 study reported that only 37.1 percent of standard blue-link citations came from the top 10 for the original query, while 36.7 percent did not rank in the top 100. Fan-out searches and other result blocks can surface different sources.
Does changing the publication date improve AI visibility?
Not by itself. Ahrefs found that cited pages were newer on average across its 2025 sample, but the effect varied substantially by platform. Update a page when facts, evidence, examples, or recommendations change, and show what changed. Cosmetic date changes do not create useful freshness.
How often should AI search visibility be tested?
For a small program, test a fixed set of commercially important prompts monthly and after material page changes. Run each prompt several times under recorded conditions. Faster monitoring may make sense for volatile news, reputation, pricing, or product-launch topics.
Share
X LinkedIn Reddit
Build Yours

Want a system
like this one?

Book a free 30-minute call. We map your situation, identify the highest-impact automation, and figure out if we are a fit.

Book Free 30-min Call