GEO and AEO: what the evidence actually supports
Two new acronyms describe getting cited by AI assistants, and most of the advice attached to them has never been tested. This page covers what the terms mean, which crawlers gate citation, and which tactics have evidence behind them. Where nobody knows, it says so.
What GEO and AEO mean, and where the terms came from
GEO (generative engine optimisation) and AEO (answer engine optimisation) mean the same thing in practice: getting content used and cited in AI-generated answers (ChatGPT, Google's AI Overviews and the like), the way SEO targets ranked results.
GEO is the term with an actual origin: a 2023 paper by Aggarwal and colleagues, which tested nine content changes and reported that the strongest raised visibility in generated answers by up to 40 per cent.
The caveat that gets dropped: it measured how a generator treats pages already in its source set, on the authors' own test rig, not on ChatGPT or Google. AEO is a marketing coinage with no origin paper; the two terms are sold interchangeably.
Underneath both sits the distinction that matters: retrieval (getting found) and citation (getting used) are separate steps. Most published GEO advice addresses the second. Site owners worry about the first.
Training, citation and live fetch are three different things
Each assistant runs several crawlers doing different jobs. Blocking the wrong one has consequences nobody intended.
OAI-SearchBot
OpenAI's search crawler: it decides whether a page can appear in ChatGPT's answers. OpenAI's documentation is unambiguous: opted-out sites "will not be shown in ChatGPT search answers", though they can still appear as navigational links.
GPTBot
Training. Blocking it keeps a site out of model training; it has no bearing on whether ChatGPT can cite the page.
ChatGPT-User
A live fetch made when a user's question sends ChatGPT to read a specific page mid-conversation.
Claude-SearchBot, PerplexityBot
The citation crawlers for Claude and Perplexity. ClaudeBot, without the suffix, is the training crawler.
Google-Extended
Neither crawling nor citation. A permission token: whether content already crawled by Googlebot may be used for Gemini training. Disallowing it does not remove a site from AI Overviews. Nothing does, short of leaving Google Search.
anthropic-ai, Claude-Web
These appear in nearly every published block-the-AI-crawlers snippet. Anthropic has never documented either.
OpenAI also notes roughly a day's propagation delay after a robots.txt change. A check made minutes after an edit proves nothing.
Whether ChatGPT runs on Bing's index
Where the claim came from
The claim is repeated everywhere and was never quite true. What existed was a privacy disclosure: OpenAI's help documentation said ChatGPT might share disassociated search queries with providers such as Bing. That named Bing as a recipient of queries, not a source index.
The wording has since changed twice: it became "sometimes partners with other search providers", and Shopify was added as a second named provider. The restructured Microsoft and OpenAI agreement of October 2025 covers compute, intellectual property and Azure exclusivity, not a search index.
What actually gates ChatGPT citation
OpenAI's actual requirement: OAI-SearchBot must be allowed to crawl the site, and the host or CDN must allow OpenAI's published address ranges. OAI-SearchBot crawls from Azure address space, because OpenAI runs on Azure. That gets misread as a Bing dependency.
Where Bing does matter
- Microsoft Copilot genuinely runs on Bing, with no separate crawler of its own.
- Bing Webmaster Tools is currently the only publisher-facing report that shows AI citations at all.
- Google's AI surfaces run on Google's index. Perplexity runs and claims its own. Anthropic has not documented what backs Claude's web search.
Being in Bing is worth having; it is not a proxy for anything else. The Bing Search APIs were switched off in August 2025, removing the piggyback route third-party products used.
Check nothing is already blocking the crawlers
Before any content advice matters, the requests have to arrive. Some sites block AI crawlers without anyone having decided to, because the control sits at the CDN.
Cloudflare separates its controls into training, search and live user fetches. An owner blocking model training will, unless careful, also switch off eligibility to be cited. Opposite intentions, adjacent switches.
The one-command check
curl -A "OAI-SearchBot" -o /dev/null -w "%{http_code}" https://example.com/
Anything but a 200 with the actual HTML means everything below is irrelevant until fixed. The same test applies to Claude-SearchBot, PerplexityBot, Googlebot and bingbot. A challenge page removes the site from that surface as effectively as a robots.txt disallow, and silently.
Also check that robots.txt as served matches robots.txt as written; some CDNs now manage or prepend to that file.
The September 2026 Cloudflare change
One dated change makes this urgent. Cloudflare's documentation deprecates the legacy Block AI bots option on 15 September 2026: from that date, "Mixed-purpose crawlers that combine Search and Training will also be blocked by all configurations to block AI training".
Cloudflare's announcement of 1 July 2026 names them: "multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training".
- Today the toggle excludes only crawlers that do both jobs, so Googlebot passes.
- From that date the same untouched setting is judged against every behaviour a crawler has, most restrictive winning.
- The fix, for anyone who switched on AI blocking once and moved on: set Search to Allow and record the mixed-purpose opt-out before the date.
Structured data is being read as text, not as structure
Adding schema markup for AI assistants is near-universal advice. Mark Williams-Cook tested it in May 2026: publish deliberately invalid JSON-LD with a fabricated address absent from the visible text, then ask the assistants about the business.
ChatGPT and Perplexity both returned the fabricated address; Perplexity said it had found it in the page's embedded structured data. The markup was reaching them as text, not as parsed structure.
One test, held lightly. But no published study demonstrates a citation benefit from schema, and this offers a mechanism for why there might not be one.
None of which argues against schema. It earns its place for Google's rich results, where the benefit is documented and mechanical. The narrower claim, that assistants parse it as structure, is the one without support.
The file this site serves that nobody has confirmed reading
llms.txt is a proposed convention: a plain text index at the site root, pointing at markdown versions of the important pages, on the theory that an assistant would rather read clean prose than parse navigation.
This site serves one, plus a markdown mirror of every content page. No AI vendor has confirmed reading it or documented support for the convention.
It is served here anyway, as a bet rather than evidence. The cost is close to zero, the mirrors help anyone reading in plain text, and if the convention is adopted the work is already done. Much published advice presents it as established practice. It is not.
Reddit is read constantly and cited rarely
Reddit's prominence in AI answers has produced a tactic: get the brand mentioned in threads, and the assistants will repeat it. The measurement does not support the tactic.
What the studies found
- Ahrefs, 1.4 million ChatGPT prompts: Reddit cited in 1.93 per cent of retrievals, while making up 67.8 per cent of the URLs retrieved and then not cited.
- Yext, 6.8 million citations from 1.6 million responses: forums at around 2 per cent.
- SE Ranking, 100,000 keywords in June 2024: Reddit out of the top ten linked domains in AI Overviews entirely.
- Semrush, 248,000 Reddit URLs: 80 per cent of cited posts had fewer than twenty upvotes. Citation does not track quality.
The rate is also unstable in a way content cannot explain: ChatGPT's Reddit citation rate fell roughly sixfold over six weeks in late 2025, with no change to Reddit's content.
The licensing detail
The figure everyone quotes for the Google deal was never confirmed by either party, and Reddit's filings name no counterparty. Google's own announcement states the arrangement "does not change Google's use of publicly available, crawlable content for indexing, training, or display".
Reddit's robots.txt is a blanket disallow with no carve-outs, Googlebot included. The access runs through contract instead.
The rules problem
Reddit prohibits paid promotion outright, with stated sanctions up to domain-level bans.
When a university research team ran language model personas in a subreddit, with ethics approval, Reddit's chief legal officer called it "deeply wrong on both a moral and legal level" and issued formal demands. A consultant doing the commercial version has no such cover. That alone is enough to advise a client against it.
The question nobody has measured
Every study showing a content change improves AI visibility starts from pages already being retrieved: the GEO paper measured sources the generator had already found, and the citation studies measure which retrieved pages get quoted.
Whether any on-page change causes a never-cited page to start being cited has not been measured by anyone. That is the step site owners actually care about, and the confident advice about it is guesswork.
One thing is mechanical: a page the citation crawlers cannot fetch cannot be cited. That is the only place where the vendors document cause and effect. Everything downstream is belief, including what I do on my own site.
What this site does, and what I cannot prove about it
This site serves:
- llms.txt and a markdown mirror of every content page
- JSON-LD on every page, including a shared entity node
- a sitemap
Nothing at the CDN blocks any AI crawler, verified by requesting pages as each crawler rather than reading the settings page.
I cannot demonstrate that any of it causes a citation. The site is new, and a handful of clicks is not a dataset. It costs almost nothing, and knowing would take a controlled test nobody has run.
The subject is being sold hard by people who do not distinguish between what they have measured and what they assume. The gap between the two is most of the published advice.
Questions this raises
Is GEO different from SEO, or is it a rebrand?
Mostly a rebrand, with one real difference. Being crawlable, answering a question directly and being mentioned elsewhere is the same work it always was. What genuinely differs: citation is gated by separate crawlers with their own permissions, so a site can rank in Google and be invisible to an assistant, or the reverse.
Should I block AI crawlers to protect my content?
A legitimate choice, and the thing to get right is which crawler. Blocking the training agents keeps content out of model training. Blocking the search agents removes eligibility to be cited, usually the opposite of what was wanted. They are separate agents with separate names, and most published snippets conflate them.
How do I tell whether an assistant is citing me?
Imperfectly. Bing Webmaster Tools reports citations across Microsoft's surfaces, currently the only publisher-facing report of its kind. Otherwise there are referral hits in analytics from assistant domains, which undercount badly because a cited page is often read without being clicked. Most of it is unmeasurable at present.