Website teams now face a practical question: which changes help people find and understand their business when search includes an AI-generated answer?
Terms such as GEO (Generative Engine Optimisation), AEO (Answer Engine Optimisation) and AI search optimisation are now used to describe techniques intended to improve how websites appear in systems such as ChatGPT, Google AI Mode and AI Overviews, Claude, Gemini, Perplexity, Grok, DeepSeek and other generative answer engines.
The quality of the advice varies as much as the terminology.
Official guidance, academic experiments and observational datasets offer different kinds of evidence. Read them on their own terms: a correlation suggests a question to investigate, while a controlled result applies within the conditions tested.
The question behind this report
What does the available evidence support, and how should a website team use it to plan improvements?
Our assessment separates website readiness from the outcomes observed in AI responses.
A website can be crawlable without being retrieved. It can be retrieved without being selected. It can be selected without being cited. It can be cited without its brand being mentioned. And it can be mentioned without generating a click or any measurable commercial outcome.
A working model of the stages
Eligible → Retrieved → Selected → Cited → Used → Brand attributed → Engaged
Treat these as distinct observations, not a sequence that every provider follows in the same order. Some internal stages cannot be directly observed from a public answer.
EXECUTIVE SUMMARY
The evidence reviewed for this report supports ten broad conclusions.
- 01Traditional SEO remains foundational. Google explicitly states that its generative AI search experiences are rooted in its core Search ranking and quality systems. Crawlability, indexing, useful content, technical quality and relevance therefore remain essential foundations.
- 02AI citation behaviour is not the same as traditional ranking. Generative systems can issue multiple related retrieval queries, select a subset of retrieved sources and cite only some of them. A page's traditional ranking can help without fully predicting whether it will be cited.
- 03Content quality and document-level structure matter more than superficial “AI wording tricks”. Peer-reviewed 2026 research found that citation visibility was influenced more strongly by interpretable document-level properties than isolated lexical edits.
- 04Clear structure, concise sections, tables, question-and-answer formats and supporting evidence are repeatedly associated with AI citation visibility. Microsoft recommends these directly, while large observational studies report positive associations.
- 05Freshness matters, but freshness alone is not sufficient. A 17-million-citation Ahrefs study found that AI assistants generally cited fresher content than conventional organic results, while newer research also shows that established older pages can continue to win citations where relevance is strong.
- 06Structured data remains useful, but its direct effect on AI citations is unproven. Google says structured data is not required for generative AI search and that there is no special AI schema. Ahrefs found no meaningful citation increase after tracking 1,885 pages that added JSON-LD against 4,000 controls.
- 07llms.txt is currently optional and weakly supported as an AI visibility tactic. Google says it neither helps nor harms visibility in Google Search, while an Ahrefs analysis of 137,210 domains found that 97% of valid llms.txt files received no requests during May 2026.
- 08Citation and brand visibility are different outcomes. In one Semrush study, 61.7% of observed brand appearances were citation-only: the source appeared as a link, but the brand was not named in the answer.
- 09Topic-level visibility is more useful than tracking one prompt. A Semrush study across 1,094 ChatGPT categories found that only 15.2% had a clear brand owner, suggesting that visibility should be assessed across groups of related prompts and intents rather than one isolated query.
- 10Accessibility is becoming relevant to AI agents as well as people. OpenAI states that ChatGPT's browser agent uses ARIA information to understand page structure and interactive elements, while Chrome describes the accessibility tree as a primary machine-readable representation for agents.
What this means for a website team
Keep a clear record of the website problem, the change made and the result observed. Check delivery separately from later visibility, and use missing or contradictory evidence to guide the next review.
1. GEO IS REAL, BUT THE EVIDENCE IS STILL DEVELOPING
The 2023 paper “GEO: Generative Engine Optimization”, later accepted at KDD 2024, introduced an experimental framework for measuring and improving visibility in generative-engine responses.
The study reported that GEO methods could improve visibility by up to 40% in its experimental environment, while also finding that the effectiveness of individual techniques varied across domains.
Read that number within the study’s conditions.
The result does not mean that adding a particular heading, schema type or writing style will universally improve AI visibility by 40%. It demonstrates that visibility in generative engines can be influenced — and that different content and query contexts respond differently.
Later research examines which content properties matter within those experimental settings.
A peer-reviewed ACL 2026 paper, “Think Before Writing: Feature-Level Multi-Objective Optimization for Generative Citation Visibility”, found that citation behaviour was more strongly influenced by document-level content properties than by isolated lexical changes. Its feature-level approach outperformed token-level rewriting approaches while maintaining or improving content quality across the tested generative engines.
For a website team, the relevant lesson concerns the quality of the whole document.
Start with the information a reader needs and organise it so the relationships are clear.
Our practical interpretation is to improve relevance, structure and substance before experimenting with surface wording.
2. AI VISIBILITY IS A PIPELINE, NOT A POSITION
Traditional search trained site owners to think in positions:
“We rank number 3 for this keyword.”
Generative systems do not expose visibility in the same way.
A useful model is:
Eligibility — Can the platform discover, crawl, index or otherwise access the content? Retrieval — Does the content appear relevant enough to enter the candidate source set? Selection — Does the system choose to inspect or use the source further? Citation — Is the source visibly attributed in the generated answer? Use / absorption — Does information from the source materially contribute to the answer? Brand attribution — Is the brand itself named or recognised? Engagement — Does the visibility lead to a click, visit, enquiry or other measurable action?
Keep the observations separate, and mark internal stages as unknown when they cannot be inspected.
Retrieval is not citation
Ahrefs analysed 1.4 million ChatGPT prompts and found that ChatGPT cited only around half of the URLs that entered the observed retrieval journey.
The study also found that semantically relevant titles and retrieval snippets were associated with citation selection.
This creates an important distinction:
A website may already be technically discoverable and retrievable, yet still lose at the selection or citation stage.
Google explicitly uses query fan-out
Google describes query fan-out as part of how its generative Search experiences work. The model can generate several related queries, retrieve additional relevant pages, and use those results to answer the user's broader question.
A user may ask:
“How do I improve my website's AI visibility?”
But the system could retrieve information around generative search optimisation, AI citations, structured data, crawler access, content freshness, brand authority, or other subtopics.
That means optimising only for the exact wording of one prompt is an incomplete strategy.
Bing now exposes grounding queries
Microsoft's AI Performance reporting in Bing Webmaster Tools exposes grounding queries: grouped phrases associated with the retrieval activity that led to a site's content being cited.
Microsoft also exposes total citations, cited pages, page-level citation activity, citation trends, query-to-page mapping, topics, intents and Citation Share.
Critically, Microsoft states that these metrics do not represent ranking, authority or importance. They are measurements of observed citation activity.
Use those metrics to describe observed activity, with the scope of the reporting visible.
3. TRADITIONAL SEO STILL MATTERS
Existing SEO work still provides the foundation for a public website.
AI measurement adds questions about the responses generated from available information.
Google's official 2026 guidance is explicit: generative AI features in Google Search remain rooted in Google's core Search ranking and quality systems.
Google describes retrieval-augmented generation as relying on its Search systems to retrieve relevant, current web pages from its index before generating responses.
This means familiar fundamentals continue to matter: crawlability, indexability, canonicalisation, useful internal linking, clear page purpose, technically accessible content, JavaScript SEO where applicable, page experience, duplicate-content management, useful original content and search-quality compliance.
Microsoft's AI Performance guidance similarly connects AI inclusion with clear structure, depth, evidence, freshness and content eligibility.
Plan AI visibility work alongside search, content and website maintenance, with separate measures for the outcomes each channel exposes.
4. RANKING IN GOOGLE IS HELPFUL, BUT IT DOES NOT GUARANTEE AN AI CITATION
Traditional search visibility and generative citation visibility overlap, but they are not identical.
This is partly explained by query fan-out.
The generative engine may retrieve sources using related searches that differ substantially from the literal user prompt. It may then select only some of those sources to support different claims within the final response.
For planning, broaden the question beyond one keyword.
A useful review asks:
“Do our important pages answer the related questions our customers ask, and what evidence shows whether providers use them?”
This is one reason single-prompt monitoring is fragile.
A brand may appear for one prompt and disappear for another closely related question.
5. CLEAR, USEFUL DOCUMENT STRUCTURE HAS SOME OF THE STRONGEST PRACTICAL SUPPORT
Clear, useful structure is a practical priority supported by several of the sources reviewed here.
Microsoft recommends descriptive headings, concise sections, tables, FAQ-style content where useful, supporting examples and data, freshness, depth and reduced ambiguity.
Semrush's 2026 content study reported positive associations between AI citations and several text characteristics, including clarity and summarisation, E-E-A-T-style signals, Q&A format, section structure and structured data elements.
These findings are observational. They should not be interpreted as proof that adding a table or FAQ will cause an AI platform to cite a page.
Read that association alongside the results of controlled academic experiments.
The ACL 2026 FeatGEO work found that interpretable document-level features were more influential than isolated lexical edits.
Practical implication
A page should make it easy to identify what the page is about, which question each section answers, what is fact versus opinion, what entities are being discussed, what evidence supports important claims, what the key comparison or conclusion is, and which information is current.
This is good writing and good information architecture first.
The possible benefit to machine interpretation follows from making the information easier to understand.
6. FAQS CAN HELP — BUT FAQ CONTENT AND FAQ SCHEMA ARE NOT THE SAME THING
FAQs are useful when a page needs to answer recurring customer questions.
The format has support in both guidance and observational research.
Microsoft explicitly recommends FAQ-style sections where appropriate because they can make information easier to understand and reference.
Semrush also found an association between Q&A formatting and cited content.
But this should not be confused with a claim that FAQ structured data is itself an AI citation factor.
Google retired FAQ rich-result support in May 2026 and subsequently removed the feature's documentation.
A practical distinction
Write useful answers where readers need them. Treat any FAQ markup as a separate implementation choice, without assuming that it will increase citations.
7. STRUCTURED DATA HELPS MACHINES UNDERSTAND CONTENT, BUT DIRECT CITATION LIFT IS UNPROVEN
Evaluate structured data by the facts and supported features it describes.
Its established uses are specific:
JSON-LD and schema.org markup can clarify organisations, people, products, events, articles, breadcrumbs, local business information and other entities.
Structured data also remains relevant to eligibility for supported rich results in conventional Google Search.
However, Google explicitly states that structured data is not required for generative AI search and that there is no special schema.org markup that websites need to add for generative AI visibility.
This distinction is reinforced by empirical research.
Ahrefs tracked 1,885 pages that added JSON-LD schema between August 2025 and March 2026 and compared them with 4,000 control pages. The study found no meaningful improvement in citations across ChatGPT, Google AI Mode or AI Overviews attributable to adding schema.
The study does not settle every use of structured data.
It does challenge the assumption that adding markup alone produces more citations.
How to use that finding
Structured data is a machine-understanding and search-feature tool. It may contribute to a cleaner information environment, but schema presence alone should not be treated as proof of AI citation readiness.
8. LLMS.TXT IS NOT CURRENTLY A MEANINGFUL UNIVERSAL GEO REQUIREMENT
Assess llms.txt as an optional publishing experiment.
The proposed file gives AI systems a concise machine-readable summary of a site's key content.
Chrome describes it as an emerging convention and its Lighthouse agentic-browsing tooling can check for it. However, a missing file is treated as Not Applicable, because it remains optional.
Google's Search guidance is clearer still: Google Search ignores llms.txt; maintaining the file neither helps nor harms Google Search visibility or rankings.
Ahrefs analysed 137,210 domains using its analytics and bot data. Of the roughly 38,000 domains in the sample with valid llms.txt files, 97% received no requests to the file in May 2026.
The observed requests describe that sample and period; they cannot predict future adoption.
Keep the effort proportionate to the evidence. A published file is neither a universal requirement nor proof that a provider used it.
9. FRESHNESS MATTERS — BUT SIMPLY CHANGING A DATE DOES NOT MAKE A PAGE USEFUL
Freshness studies offer a useful observation to investigate, rather than a schedule for rewriting pages.
Ahrefs analysed nearly 17 million cited URLs across seven AI/search surfaces and found that AI assistants generally cited fresher pages than conventional organic Google results.
Ahrefs reported a 25.7% lower average publication age for URLs cited by the AI assistants it grouped, compared with organic results. That association does not establish that changing a publication date earns citations.
However, the newer 1.4-million-prompt ChatGPT study adds an important qualification.
The average cited page in that dataset was not brand new. Established pages could continue to be selected when their retrieval information and relevance remained strong.
What to update
Maintain content because facts, products, regulations, statistics or recommendations have changed — not because changing the date field is an optimisation trick.
Freshness should mean continued accuracy.
10. CITATION AND BRAND VISIBILITY ARE SEPARATE OUTCOMES
Being used as a source does not guarantee that users become aware of the source.
Semrush's 2026 “ghost citations” study found that 61.7% of observed appearances were citations where the brand itself was not mentioned in the answer.
The same study found substantial differences between AI engines.
This gives us at least three distinct visibility outcomes:
- 01Citation only — the page appears as a source, but the brand is not named.
- 02Citation + mention — the page appears as a source and the brand is explicitly named.
- 03Mention without citation — the brand appears in the answer without an attributable source link.
These are commercially different.
A citation may strengthen source attribution and potentially generate referral traffic.
A brand mention may increase awareness even when it produces no direct visit.
A measurement platform should therefore avoid treating all appearances as equivalent.
Generative answers can vary between prompts, models, sessions and time periods.
That makes “we appeared for this one prompt” a weak success metric.
A Semrush 2026 study tracked 1,094 topic categories in ChatGPT from January to June. Each category contained five representative prompts spanning different stages of user intent.
Only 15.2% of categories had a clear brand owner according to the study's definition.
More than half were unsettled.
The study also found that a meaningful lead was more stable month to month than a narrow advantage.
How to apply the finding
Our recommendation is to review a stable set of related customer questions, with provider coverage and dates recorded for each one.
For example, a business software company might monitor a group containing “What is custom business software?”, “Best custom software development companies”, “Custom software vs SaaS”, “When should a business build its own platform?” and “How much does custom software cost?”
A useful system would measure presence across the cluster, not celebrate one favourable response.
12. BUILD A CREDIBLE BUSINESS PRESENCE
A business’s reputation extends beyond the pages it controls.
The evidence does not reveal a universal weighting for backlinks, mentions or other reputation signals.
Generative systems can draw on search-index signals, entity relationships, first-party websites, independent publications, forums, videos, product data, knowledge graphs and other retrieval sources.
The mixture depends on engine, query and topic.
Google explicitly warns against seeking inauthentic mentions as an AI optimisation tactic, noting that its generative features remain subject to quality and anti-spam systems.
At the same time, external recognition remains strategically important because brands do not exist only on their own websites.
What to build
Check the result, then decide what comes next
Keep your own website accurate and build legitimate references through the work you do. Which sources appear in an answer depends on the provider, topic and question.
This is especially relevant for new brands.
A technically excellent one-month-old website cannot instantly manufacture years of reputation, references, independent reviews or topic association.
What it can do is build those signals honestly.
13. ACCESSIBILITY IS BECOMING PART OF AGENT READINESS
Accessibility helps people understand and use a website.
That human purpose remains the reason to prioritise it.
Some browser agents also use the same semantic information to interpret an interface.
OpenAI states that ChatGPT's browser agent uses ARIA roles, labels and states to understand page structure and interactive controls.
Chrome's 2026 guidance says agents can perceive a site through three primary representations: screenshots, raw HTML / DOM, and the accessibility tree.
Chrome describes the accessibility tree as a high-value semantic representation containing roles, names and states that agents can use to understand interactive functionality.
In June 2026, Chrome introduced an Agentic Browsing category in Lighthouse. Its audits include checks around accessibility, interaction stability and emerging WebMCP integration.
The important qualification is that this should not be misrepresented as proof that ARIA improves AI search citations.
What the guidance supports
Accessible, semantically clear interfaces are easier for browser agents to understand and operate.
That is agent readiness, not a proven citation ranking factor.
The web now has two increasingly important non-human audiences.
Retrieval systems search, index, retrieve, rank, ground and cite information. Their concerns include crawlability, relevance, quality, freshness, source selection, authority and citation.
Action-oriented agents navigate interfaces and perform tasks. Their concerns include accessibility trees, programmatic names, stable layouts, deterministic controls, forms, semantic roles and emerging protocols such as WebMCP.
Chrome explicitly distinguishes these two stages: agents searching the web and agents using the web.
Test the two tasks separately.
A website could be excellent at getting cited and poor at allowing an agent to complete a booking.
Or it could be highly agent-friendly but almost invisible in retrieval systems.
These should not be collapsed into one “AI-ready” label.
15. USE THE MEASUREMENT EACH PLATFORM EXPOSES
One of the clearest developments in 2026 is that major platforms are beginning to expose real AI visibility data.
Microsoft
Bing Webmaster Tools' AI Performance reporting now includes total citations, average cited pages, cited URLs, grounding queries, page-to-query mapping, visibility trends, topics, intents, Citation Share and exportable time-series data.
Microsoft repeatedly cautions that these metrics are observational.
A rise in citations after a content change does not automatically prove that the change caused the rise.
In June 2026, Google launched dedicated Generative AI performance reports in Search Console for a subset of websites.
The reports expose visibility within AI features such as AI Overviews and AI Mode, including generative-AI impression data and related dimensions.
OpenAI
OpenAI states that publishers can allow OAI-SearchBot so public content can be discovered, surfaced and cited in ChatGPT Search. Referral links from ChatGPT include utm_source=chatgpt.com, allowing publishers to measure inbound traffic through analytics.
A practical direction for measurement
Record the observations the provider exposes, with dates and coverage. Check the actual result before deciding whether a website change helped.
16. AN EVIDENCE MATRIX FOR GEO TACTICS IN 2026
17. WHAT WE CAN SAY WITH CONFIDENCE
AI visibility can be influenced.
Academic and industry research shows that document characteristics and website quality can affect generative citation visibility.
AI visibility is engine-specific.
Different engines have different retrieval systems, source mixes, presentation formats and citation behaviour. Optimising for “AI” as though it were one search engine is an oversimplification.
Technical readiness is necessary, but not sufficient.
A page must generally be accessible to the relevant retrieval system before it can be used. But being crawlable does not mean it will be retrieved, and retrieval does not guarantee citation.
Content quality still matters.
Generative systems have not removed the need for original, useful, relevant information. If anything, systems that synthesise answers from multiple sources increase the value of content that contributes something specific enough to deserve retrieval.
Structure helps.
Clear information architecture, descriptive headings and logically organised sections make documents easier for both humans and machines to interpret.
Authority cannot be installed with a snippet.
Technical optimisation can improve a website's readiness. It cannot instantly create reputation, years of expertise, independent coverage, customer experience, citations from authoritative publications or broad brand recognition. Those remain earned signals.
Measurement must be continuous.
AI models, retrieval systems, citation interfaces and source preferences change. A tactic that appears effective today should be tested against observable outcomes rather than assumed to remain effective indefinitely.
18. QUESTIONS THE EVIDENCE DOES NOT SETTLE
Make uncertainty part of the decision, rather than filling it with a universal rule.
As of August 2026, the evidence does not justify confidently claiming that one schema type will cause ChatGPT to cite a page; llms.txt materially improves universal AI visibility; adding FAQs automatically boosts citations; every AI engine rewards the same signals; a particular keyword density improves generative selection; every third-party mention improves AI authority; one score can reliably predict all AI visibility; ranking first in Google guarantees citation; a citation guarantees brand awareness; a brand mention guarantees referral traffic; or a traffic increase after a GEO change proves causation.
A measurable observation is a useful starting point.
Explaining its cause requires further evidence.
19. A MORE USEFUL MODEL FOR EVALUATING AI VISIBILITY
We propose separating AI visibility into seven layers.
Layer 1 — Eligibility
Can the relevant platform access the page? Typical checks: crawler permissions, robots directives, indexability, HTTP status, canonical behaviour, render accessibility, sitemap and discovery pathways.
Layer 2 — Retrieval readiness
Is the page likely to be understood as relevant to the topic or grounding query? Typical considerations: clear title and page purpose, semantic relevance, useful body content, topical coverage, internal contextual links, current and accurate information.
Layer 3 — Information quality
Does the page contain enough useful information to contribute to an answer? Typical considerations: first-hand or expert information, data, definitions, comparisons, examples, methodology, limitations, supporting sources.
Layer 4 — Authority and entity clarity
Can systems confidently understand who produced the information and how the organisation, person, product or topic relates to the wider web? Typical considerations: consistent organisation/entity details, credible authorship, external references, real-world reputation, coherent topic focus.
Layer 5 — Citation visibility
Is the content actually appearing as an attributable source? Measures could include citations, cited pages, citation frequency, Citation Share, grounding-query coverage and engine-specific visibility.
Layer 6 — Brand attribution
When the site's information is used, is the brand actually named? Measures could include brand mentions, citation + mention combinations, ghost citations, sentiment and contextual association.
Layer 7 — Engagement
Does AI visibility produce measurable user behaviour? Measures could include AI referral visits, landing pages, conversions, assisted conversions, enquiries and revenue where attribution is possible.
This model deliberately separates readiness from outcomes.
It also makes missing evidence visible, so a team can decide what to investigate next.
20. IMPLICATIONS FOR WEBSITE OWNERS
Use the research to organise practical website work:
- 01Make sure important content is accessible to the retrieval systems you want to appear in.
- 02Fix ordinary technical SEO problems before inventing exotic AI-specific ones.
- 03Create genuinely useful content around the topics customers actually care about.
- 04Structure that content clearly enough that humans and machines can identify its key information.
- 05Support important claims with evidence, examples and sources.
- 06Keep time-sensitive information substantively current.
- 07Use structured data accurately, but do not treat it as a magic citation switch.
- 08Treat llms.txt and emerging agent protocols as experiments, not established ranking factors.
- 09Improve accessibility because it benefits people and increasingly helps agents understand interfaces.
- 10Measure AI citations, mentions, impressions and referrals independently.
- 11Monitor topics across multiple related prompts rather than celebrating one favourable AI answer.
- 12Re-test assumptions as platforms and models change.
21. WHAT THIS MEANS FOR ANSWERMETRIX
AnswerMetrix brings website findings, proposed changes and recorded observations into a workflow for ongoing improvement.
The product should help a team distinguish the condition of a page from its appearance in an AI answer.
A readiness score can be useful when it describes observable technical, structural and content conditions.
It becomes misleading if presented as proof that a website will be cited.
Inspect saved mention evidence, available response excerpts and returned sources. Review connected search and analytics data separately. The wider measurement model in this report includes external capabilities that AnswerMetrix does not currently integrate.
The research reviewed in this report supports an iterative approach:
Discover → Review → Approve → Deliver → Verify → Measure → Learn
A deployed optimisation should not be assumed to work merely because the implementation succeeded.
The next question
“What changed on the public page, and what do later measurements show?”
And even then, correlation should not automatically be presented as causation.
Keeping the work and the observations connected gives a team a clearer basis for its next decision.
22. METHODOLOGY
This report reviews published research, commercial studies and official platform guidance. It contains no original AnswerMetrix crawl dataset.
Every dataset discussed belongs to the named researchers or publishers. The report’s evidence ratings and practical recommendations are AnswerMetrix’s editorial assessment.
The report reviews and compares official platform documentation from Google Search, Microsoft Bing, OpenAI and Chrome/web.dev; peer-reviewed and academic research including the original GEO research framework and 2026 ACL research into feature-level generative citation optimisation; large commercial datasets from Ahrefs and Semrush; and current platform measurement capabilities from Google Search Console, Bing Webmaster Tools and OpenAI.
Evidence was weighted qualitatively.
Strong — Supported by official platform documentation, reproducible mechanisms, multiple high-quality datasets, or strong academic evidence.
Moderate — Supported by repeated observational evidence or credible platform guidance, but causation is not established.
Emerging — Technically plausible or newly supported, but too new or insufficiently studied to treat as settled.
Weak / unproven — Popularly recommended but lacking strong direct evidence.
Contradicted / platform-specific — A tactic may be unsupported by a particular engine, explicitly rejected by a platform, or only applicable in limited environments.
23. LIMITATIONS
This field changes rapidly.
Platform behaviour can change. Models, indexes, grounding providers, user interfaces and citation policies can change without notice. Findings should therefore be treated as a snapshot of the evidence available through 27 August 2026.
Commercial datasets are observational. Ahrefs and Semrush provide unusually large datasets, but they are commercial organisations measuring platforms from outside. Correlations found in these datasets do not automatically reveal proprietary ranking or citation mechanisms.
Platform data is incomplete. Even first-party dashboards can be aggregated or sampled. Microsoft explicitly notes that its AI Performance reports are not complete logs of every citation instance.
Engines are not interchangeable. A result observed in ChatGPT should not automatically be generalised to Google AI Mode, Gemini, Claude, Perplexity, Grok, DeepSeek or other engines.
Visibility is not business value. A citation, impression or brand mention is an intermediate outcome. Organisations should ultimately connect AI visibility to awareness, qualified visits, leads, conversions, customer acquisition or other meaningful objectives.
24. PUT THE EVIDENCE TO WORK
Start with a website that serves its customers well: accessible pages, clear information, accurate facts and working journeys.
Then inspect how each relevant platform represents it.
A generated answer can involve retrieval, source selection, synthesis and attribution. What you can observe depends on the provider and the evidence it returns.
Search checks help establish whether important pages can be found.
Content review helps establish whether those pages answer useful questions.
Verified facts and credible sources help readers assess the information.
Accessible interfaces help people complete the next step.
AI response measurements add another view: what was mentioned or cited in the answers sampled, at the time they were recorded.
A useful review cycle
Choose a worthwhile change and record the reason for it.
Check the result, then decide what comes next
Verify the public page, compare later observations and investigate differences. Keep the query set, dates and provider coverage visible so the comparison remains meaningful.
Progress should be explainable.
A team should be able to say what it changed, what it checked and what it still does not know.
That makes the next investment in content, implementation or measurement easier to assess.
The value lies in better decisions and useful website work, supported by evidence that readers can inspect.
REFERENCES
Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K. & Deshpande, A. “GEO: Generative Engine Optimization.” arXiv:2311.09735. https://arxiv.org/abs/2311.09735
Liu, Z. & Xu, P. “Think Before Writing: Feature-Level Multi-Objective Optimization for Generative Citation Visibility.” Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics, 2026. https://aclanthology.org/2026.acl-long.929/
Google Search Central. “Optimizing your website for generative AI features on Google Search.” 2026. https://developers.google.com/search/docs/fundamentals/ai-optimization-guide
Google Search Central. “Introducing Search Generative AI performance reports in Search Console.” 3 June 2026. https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports
Google Search Central. “Latest Google Search Documentation Updates.” 2026. https://developers.google.com/search/updates
Microsoft Bing Webmaster Tools. “AI Performance in Bing Webmaster Tools.” 2026. https://www.bing.com/webmasters/help/ai-performance-9f8e7d6c
OpenAI Help Center. “Publishers and Developers — FAQ.” 2026. https://help.openai.com/en/articles/12627856-publishers-and-developers-faq
Chrome for Developers. “A developer toolkit to make your website agent-ready.” 22 June 2026. https://developer.chrome.com/blog/agent-ready-toolkit
web.dev. “Build agent-friendly websites.” 2026. https://web.dev/articles/ai-agent-site-ux
Chrome for Developers. “llms.txt — Lighthouse Agentic Browsing.” Updated 5 May 2026. https://developer.chrome.com/docs/lighthouse/agentic-browsing/llms-txt
Ahrefs. “Why ChatGPT Cites One Page Over Another (Study of 1.4M Prompts).” 15 April 2026. https://ahrefs.com/blog/why-chatgpt-cites-pages/
Ahrefs. “AI Assistants Prefer to Cite ‘Fresher’ Content (17 Million Citations Analyzed).” 28 July 2025. https://ahrefs.com/blog/do-ai-assistants-prefer-to-cite-fresh-content/
Ahrefs. “We Tracked 1,885 Pages Adding Schema. AI Citations Barely Moved.” 2026. https://ahrefs.com/blog/schema-ai-citations/
Ahrefs. “We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read.” 15 June 2026. https://ahrefs.com/blog/llmstxt-study/
Semrush. “How We Built a Content Optimization Tool for AI Search [Study].” 14 January 2026. https://www.semrush.com/blog/content-optimization-ai-search-study/
Semrush. “Why 62% of AI citations don’t lead to brand mentions [Study].” 9 June 2026. https://www.semrush.com/blog/the-ghost-citations-study/
Semrush. “AI visibility is a topic-level game: A study of 50,000 brands in ChatGPT.” 20 July 2026. https://www.semrush.com/blog/chatgpt-topic-authority-study/
SUGGESTED CITATION
AnswerMetrix Research. (2026). The State of GEO & AI Search Visibility 2026: A research review of website discovery, selection, citation and brand representation in AI answers. AnswerMetrix.
ABOUT ANSWERMETRIX
AnswerMetrix helps teams find useful website improvements, review proposed changes, deliver supported fixes and follow recorded search and AI observations.
For more information, visit answermetrix.com.
Apply the evidence with the AI visibility optimization field guide →