AIEO: AI Engine Optimisation
What gets a page retrieved and cited in an AI answer, according to the companies that run the engines, and the boundary this site keeps.
The answer, in one paragraph. AI engine optimisation (AIEO) is the work of getting a page retrieved and cited when an AI system answers a question: Google's AI Overviews and AI Mode, Microsoft Copilot and Bing, ChatGPT search, Perplexity, Claude. The same work is also called answer engine optimisation (AEO) and generative engine optimisation (GEO). Google and Bing document the same foundation: their AI answers draw on the search index, so a page has to be crawlable, indexed, clear and useful before anything else counts. OpenAI, Perplexity and Anthropic each name a search crawler and say what blocking it does to a site's place in their answers. Google says structured data is not required for its AI features and that Google Search ignores llms.txt. Bing says structured data may support grounding and guarantees nothing. None of the pages quoted below describes a way to be cited that skips these steps.
What the engines document
One row per engine, quoted from its own documentation, fetched on 2026-10-09.
| Engine | What its documentation says | Source |
|---|---|---|
| Google AI Overviews and AI Mode | The features are rooted in our core Search ranking and quality systems. To be eligible, a page must be indexed and eligible to be shown in Google Search with a snippet. AEO and GEO are still SEO. | Google Search Central, Optimizing your website for generative AI features, updated 2026-07-10 |
| Bing and Microsoft Copilot | Bing and Copilot search experiences rely on the same core crawling, indexing, and ranking foundation as traditional search.Bing lists eligibility for Grounding results and citationsamong the aims of the guidelines. | Bing Webmaster Guidelines |
| ChatGPT search | Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links. | OpenAI, Overview of OpenAI crawlers |
| Perplexity | PerplexityBot is designed to surface and link websites in search results on Perplexity.Perplexity recommends allowing it in robots.txt. | Perplexity, Perplexity crawlers |
| Claude | Blocking Claude-SearchBot prevents our system from indexing your content for search optimization, which may reduce your site's visibility and accuracy in user search results. | Anthropic, crawler help page, modified 2026-04-07 |
Google, in more detail
Google published its guide to generative AI search on 2026-05-15 and updated it on 2026-07-10. In the guide, Google names two techniques behind the answers: retrieval-augmented generation over the Search index, and query fan-out, a set of related queries the model generates to fetch more results.
- What helps. Unique, first-hand content that goes beyond common knowledge. In Google's words,
Commodity content (for example, something like "7 Tips for First-Time Homebuyers") is often based on common knowledge
and adds little. Pages organised in sections with clear headings, crawlable, and with a good page experience, which includesreducing latency
. - A Search Console control. A site must also be included in generative AI features through a Search Console setting. Include is
the default control for all properties
(Search generative AI control). - What Google Search ignores. llms.txt and other AI text files: creating them
will neither harm nor help your site's visibility or rankings in Google Search, as Google Search ignores them
. Chunking content into small pieces, rewriting content for AI systems, and inauthentic mentions are on the same list. - Structured data.
Structured data isn't required for generative AI search, and there's no special schema.org markup you need to add.
Google still recommends it for rich results. - Tools.
No third-party tool has access to our internal ranking or AI systems.
Bing, in more detail
The Bing Webmaster Guidelines cover Bing search, Copilot and grounding API results in one document. The points that bear on citation:
- Discovery. IndexNow, XML sitemaps, crawlable internal links and links from relevant sites. Sitemaps should
List only canonical URLs
. - One URL per piece of content.
Duplicate URLs dilute signals and reduce Bing's confidence in selecting a URL for grounding results or citations.
- Rendering. Avoid
Hiding critical content behind client-side rendering
andExcessive or unnecessary HTTP requests to render the content
. - Content. Content should
Be easy to understand without external context
, facts and definitions should be explicit, and entities named the same way every time:Use clear and consistent naming for people, organizations, products, and locations.
- Answer first.
Place essential information near the top of the URL
, one topic per URL. - Structured data.
Structured data may support clearer grounding but does not guarantee visibility or grounding traffic. Markup must accurately reflect visible content.
The boundary this site keeps
Every item below is a practice this site does not use. Each one is named as abuse or spam in at least one engine's own policy.
| This site does not | Policy that names it |
|---|---|
| Publish a claim it cannot back. Research pages carry their receipts; the claims ledger lists what would refute each claim; docs imported from the knowledge vault are labelled as not reviewed against their sources. | Google: People-first content means content that's created primarily for people, and not to manipulate search engine rankings.(helpful content) |
| Leave duplicate URLs unmarked. Every page has one canonical link; for the three pieces cross-posted to Substack, it points at the Substack copy. | Bing: Use canonical URLs, parameter controls, and consistent URL structures to consolidate signals and improve grounding visibility. |
| Seed posts or buy mentions. | Bing names artificial social promotion schemes that simulate popularity. Google: seeking inauthentic "mentions" across the web isn't as helpful as it might seem. |
| Buy domains or build sites to own a phrase. | Google lists Creating multiple sites with the intent of hiding the scaled nature of the contentunder scaled content abuse; its site reputation policy covers third-party content placed on a host site mainly because of that host's already-established ranking signals(spam policies, updated 2026-08-28). |
| Hide text, stuff keywords, or place text aimed at language models. | Google: Hidden text or link abuse is the practice of placing content on a page in a way solely to manipulate search engines and not to be easily viewable by human visitors.Bing lists Prompt Injection and AI Manipulation. |
| Mark up anything a reader cannot see on the page. | Google: Don't mark up content that is not visible to readers of the page.(structured data guidelines, updated 2026-07-10) |
In this section
- How this site is built for AI search: each lever, what this site does, which engine documents reading it, and where to check it.
- Measurement: the Search Console report, the Bing report and an API probe, what each one counts, and the status of each on this site.
Sources
All pages fetched on 2026-10-09. Quotes are copied from the fetched text.
- Google Search Central, Optimizing your website for generative AI features on Google Search, published 2026-05-15 (announcement), updated 2026-07-10.
- Google Search Console Help, Search generative AI control.
- Google Search Central, Spam policies for Google web search, updated 2026-08-28.
- Google Search Central, General structured data guidelines, updated 2026-07-10.
- Google Search Central, Creating helpful, reliable, people-first content.
- Microsoft, Bing Webmaster Guidelines.
- OpenAI, Overview of OpenAI crawlers.
- Perplexity, Perplexity crawlers.
- Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler?, modified 2026-04-07.
Changelog
- v1, 2026-10-09: first version.