Back to blog
NO.
017
DATE
Updated 2026-07-24
READ
~17 min
KIND
Notes
STATUS
Reviewed

TAGS: SEO

GEO: Being Cited Correctly by Answer Engines

GEO is not keyword stuffing for AI search. It is making content easy for answer engines to understand, cite and verify — with 2026 data and two corrections.

Several strands of GEO discussion look like traffic arguments but are really one question: when people stop clicking through and instead read what ChatGPT, Perplexity, Google AI Overviews, Doubao or Kimi have assembled for them, how does your content get understood, cited, and paraphrased?

This is not SEO under a new name, and it is not keyword stuffing aimed at ChatGPT.

I prefer to read GEO as answer engine visibility: making content easy to identify correctly, extract correctly, and cite correctly — and making the evidence traceable, for a reader or a model.

This post covers the fundamentals and what is actually actionable. It is not a platform-by-platform case study.

Update 2026-07-02: added Google's official guide and four 2026 studies, and corrected two claims from the first version that did not hold up — I had overstated the role of both llms.txt and structured data. What changed and why is in the update log at the end.

Conclusions first

  • GEO is not "SEO for AI." The core is content that is identifiable, extractable, citable, and verifiable — not guessing which words an AI likes.
  • Google's official position: optimizing for AI Overviews / AI Mode is still SEO. Content has to be crawlable, indexable, and eligible for snippets; the AI features are built on the same index and the same quality systems.
  • The sources AI cites are no longer the Top 10 for the original query. Ahrefs, March 2026: only about 38% of AI Overview citations come from the top 10 result blocks for the same query. A year earlier that figure was about 76%.
  • Being cited is not the same as being seen. Semrush, June 2026: 61.7% of AI citations are "ghost citations" — the page is used as a source while the brand name never appears in the answer text.
  • Neither llms.txt nor JSON-LD is an "AI citation button." Google states the first is unnecessary and has no effect; the second produced no clear citation lift in Ahrefs' controlled study.

The table below collects the numbers behind this post's key claims. It serves two purposes: you can check the strength and recency of each piece of evidence without reading the whole post, and it demonstrates the "table with concrete numbers" point from the writing framework below — for an answer engine, a table like this is far easier to extract intact than a paragraph.

FactFigureSource and date
Overlap between AI Overview citations and the query's Top 10 blocks37.9% (about 76% in 2025)Ahrefs, 2026-03
"Ghost citations": cited as a source, brand absent from the answer61.7%Semrush, 2026-06
Change in AI citations after adding JSON-LD (1,885-page controlled study)No clear positive effectAhrefs, 2026-05
Effect of llms.txt on Google Search / AI featuresNot needed, no effectGoogle official guide, 2026-05
Analysis sample: citations / models / countries3.25B / 7 / 14Profound, 2026-04

What GEO actually means

GEO usually expands to Generative Engine Optimization — content optimization aimed at generative answer engines.

The term invites a misreading as "SEO for AI," and once you read it that way the tactics slide back into old habits:

  • Guess which keywords the AI likes.
  • Mass-produce near-identical Q&A pages.
  • Write headlines as inflated conclusions.
  • Wrap empty content in authoritative-sounding sentences.

These may briefly make content look more like an answer. They do not make it more trustworthy. What an answer engine needs is not more splice-able sentences but clearer entities, more stable facts, more citable evidence, and explicit boundaries.

Worth noting: Google's official guide treats AEO and GEO as industry-external vocabulary — from Google Search's point of view, optimizing for generative AI search experiences is still SEO. That framing is itself the answer: there is no separate mystical stack to invent for AI search.

So the core question is not "how do I get the AI to mention me." It is:

When an answer engine needs to answer a question, can it recognize my content as a reliable source? Can it tell precisely what I am claiming, what it applies to, what its limits are, and where the evidence sits?

SEO and GEO optimize for different things

The traditional SEO chain: the search engine discovers a page, understands it, ranks it, and the user clicks through.

GEO adds a layer. The answer engine first compresses several sources into one answer, then offers a few of them as citations or further reading. The user may never click, or may click only to verify.

That shifts where content effort pays off.

DimensionTraditional SEO cares aboutGEO cares about
Entry pointRanking on the results pageVisibility and citability inside the answer
Unit of contentPage, title, keyword, internal linkEntity, paragraph, step, conclusion, evidence
Success signalImpressions, rank, clicks, conversionBeing cited correctly, paraphrased accurately, used as evidence
RiskNot indexed, low rank, few clicksMisread, quoted out of context, buried under identical content
Writing focusMatching search intentStating the question, the conclusion, the conditions, and the boundaries

SEO does not stop mattering. Without a crawlable, indexable, clearly structured page there is no GEO to speak of. But GEO is not SEO plus a handful of AI keywords; it asks the content itself to behave like a verifiable node of information.

The 2026 data exposes a further split: being cited and being seen are also two different things. That deserves its own section, below.

How answer engines actually choose content

To be cited correctly, you first need to know how an answer gets produced.

Most AI search today runs on RAG (Retrieval-Augmented Generation), which splits in two:

  1. Retrieval: after your question, the model pulls a batch of relevant content from whatever web it can reach — news, official sites, documentation, a forum answer.
  2. Generation: the model reads that batch, extracts the key facts, and reorganizes them into an answer in its own words.

Google confirms it works this way in the official guide: AI Overviews / AI Mode retrieve relevant, reasonably fresh pages from the Google Search index and use them to ground the response. The guide also names a mechanism most people miss — query fan-out: your original question is decomposed into several related sub-queries, and Google looks for supporting sources in the results for those. "Too many weeds in my lawn" might fan out into "best lawn weed killer," "killing weeds without chemicals," "how to prevent lawn weeds."

The direct consequence: you do not have to rank in the Top 10 for the original query to be cited — and ranking in the Top 10 does not guarantee citation. Ahrefs' March 2026 update (863K SERPs, ~4M AI Overview citations) quantified it: only 37.9% of cited URLs also appeared in the top 10 result blocks for the same query, and 31% were not even in the top 100 — against roughly 76% overlap in their earlier study. In one year, "rank well and you get cited" loosened that far.

There is also a hard logical consequence in those two steps: if your content is not retrieved in step one, or is written too messily to be extracted in step two, your point never reaches the answer. GEO is therefore the inverse exercise — write and publish the way retrieval and extraction actually work.

Think of the filter as a three-layer funnel:

  • Visible (retrievable): the physical threshold. Whether the page can be crawled, how fast it loads, whether you have blocked AI crawlers — that decides whether you are in the candidate set at all. In the fan-out era this layer also requires covering the related sub-questions of a user's journey, not betting on a single keyword.
  • Legible (extractable): once fetched, the model has to understand it. It favors clear definitions, visible structure, and high information density; it struggles with emotive filler, sprawling paragraphs, and adjective-only marketing copy.
  • Credible (verifiable): to avoid making things up, models lean toward information that appears grounded and can be cross-checked against several sources. You saying you are good, alone, carries little weight.

If GEO has to fit in one line:

GEO = SEO (so the answer engine can find you) + RAG (so the answer engine is willing and able to cite you).

SEO solves the first layer. The other two are the writing and evidence problems GEO actually has to fix.

Cited is not the same as seen

Semrush's June 2026 study (3,981 domain appearances, 115 prompts, 14 countries, 4 engines) contributes an important concept: the ghost citation — the page is hung in the source list while the brand name never appears in the answer text. A reader can finish the answer without ever registering that you exist.

Their distribution:

TypeShare
Ghost citation: cited but not mentioned in the text61.7%
Both cited and mentioned13.2%
Mentioned in the text with no source link25.1%

The engine differences matter more. ChatGPT cites 87% of the time but mentions brands in the text only 20.7% — behaving like footnoted research material. Gemini is the mirror image: 83.7% mentions against 21.4% citations, more like a conversation that says brand names out loud. The same content is "visible" in completely different ways depending on the engine.

The lesson for writers is concrete: if your page answers the question but never states plainly who you are and what you make, it degrades into an anonymous reference library — used, not remembered. That is why "explicit entities" in the framework below carries more weight in 2026 than it used to.

Get the technical floor right, once

The first funnel layer is not about prose. It is a handful of basic technical settings that pay off for years once done right. The 2026 evidence also brings two popular claims in this section back to earth, so I have rewritten it to match what I now believe.

Do not accidentally block AI crawlers in robots.txt. Plenty of templates ship User-agent: * / Disallow: /, which shuts out every crawler including answer engines. For GEO, at minimum allow the major AI crawlers (GPTBot, PerplexityBot, ClaudeBot and friends) and block only the traffic you genuinely do not want. Note that Google's AI features still ride normal Google Search crawling and indexing — meeting the technical requirements, being indexed, and being snippet-eligible are the prerequisites for appearing in AI Overviews / AI Mode.

Structured data (Schema.org / JSON-LD): do it for traditional search, not as an AI citation button. My first version said it "clearly raises the probability of accurate citation." That claim needs correcting. Ahrefs' May 2026 controlled study tracked 1,885 pages that added JSON-LD against 4,000 control pages, using difference-in-differences to strip out platform trends. The result: no clear citation lift on any platform — Google AI Overviews −4.6% (small, and both groups were already declining, so not attributable to schema), AI Mode +2.4%, ChatGPT +2.2% (neither statistically distinguishable from zero). They also cite a searchVIU experiment finding that when ChatGPT, Claude, and Perplexity fetch a page live, they read the visible HTML only; hidden markup like JSON-LD is ignored. So: keep doing schema — it still earns rich results, entity understanding, and traditional SEO value — but do not expect a block of JSON-LD to buy AI citations. Definitions, tables, and numbers in the visible body outrank any hidden markup.

llms.txt: no effect on Google, a cheap option elsewhere. This is the other correction. I originally wrote that you gain by having it and lose by not. Google's official guide puts it on the mythbusting list: for Google Search / AI Overviews / AI Mode, llms.txt is not needed, and its presence or absence affects visibility neither way (the same list also covers: no special AI markup needed, no need to convert pages to Markdown, no need to manually chunk content). Vendors report otherwise — generative.qa's benchmark associates a well-structured llms.txt with a 12% higher citation rate and more consistent consumption by Perplexity and Claude — but that is vendor correlation, not causation, and it does not apply to Google. My current practice: keep the file (near-zero cost, may serve engines other than Google), stop listing it as a GEO tactic, and do no maintenance for it.

This blog is built to that standard: robots.txt blocks only the admin and CMS paths and stays open to crawlers including AI ones; every post emits BlogPosting structured data (for traditional search, not for AI citation); and an llms.txt sits at the root — understood within the boundaries above.

A more practical GEO writing framework

Above the technical floor, whether content is legible and credible comes down to how it is written. This is the checklist I would run over an article meant to be cited accurately.

The list comes from ordinary writing sense, and the 2026 industry data broadly points the same way: generative.qa's benchmark (10K prompts, 6 engines) found that pages carrying all five of an opening entity definition, self-contained answer paragraphs, tables with concrete numbers, original data with a stated methodology, and a real author byline were about 2.4× as likely to be cited as pages with none of them (vendor benchmark, correlation not causation — read the direction, not the multiplier).

1. State the question first

Open by saying what the article answers and what it does not.

Not "what is GEO," but "how does GEO differ from SEO, and what should a writer change." The clearer the question, the easier it is for an answer engine to place the article in the right context.

2. State the conclusion first

Put the core judgment up front. Do not make readers or models hunt for it inside a narrative.

This does not mean writing like a manual. It means each key claim needs a sentence carrying it. For example:

GEO is not stuffing keywords for AI search; it is making content easy for answer engines to understand, cite, paraphrase, and verify.

A sentence like that can be quoted accurately.

3. Name entities explicitly

People, companies, products, concepts, dates, places, numbers — be specific.

Avoid a chain of "this tool," "that company," "the post I mentioned." Human readers follow context; models lose the relationships during extraction.

The ghost citation data adds weight here: 61.7% of citations never put the brand into the answer. If your own page does not state who you are, what you do, and how that relates to the topic, an answer engine has even less reason to say it for you.

4. Make steps extractable

If the article carries a method, give the steps a stable structure.

Not to please a machine, but to stop the method from deforming when it is paraphrased. "Question first, conclusion first, explicit entities, extractable steps, citable evidence, stated boundaries, internal links" is itself an extractable structure.

5. Make evidence citable

When you cite something external, do not just drop a link. Say:

  • What the link points at (real anchor text).
  • Which part of it you used.
  • Which parts you did not verify.
  • Where your own judgment starts.

That is owed to the reader and to the answer engine. Otherwise a model will happily blend source material with author commentary.

6. State the boundaries

Over-generalization is GEO's biggest failure mode.

If it is experience, call it experience; if it is a guess, call it inference; if you only read the headline and abstract, do not review the paper. The clearer the boundary, the harder the content is to paraphrase wrongly.

That is why every vendor benchmark in this post is labeled "correlation, not causation" — stating the strength of your evidence is itself part of being credible.

One article need not answer everything, but it should connect the surrounding context.

SEO fundamentals, content structure, technical indexing, case retrospectives, tooling — internal links let them support each other. For a reader that is a path; for an answer engine it is a topic graph.

Language is an independent variable

If your readers are not all English speakers — cross-border B2B, for instance — one more 2026 finding matters: the language of the query reshapes which sources the AI cites.

Profound's April 2026 study analyzed 3.25B citations (7 models, 14 countries, native-language prompts only). The headline: the citation distribution you see from English prompts does not represent other language markets. Reddit is a dominant source in English, but in Japanese, Spanish, Portuguese, and Arabic markets the supply of citable content is structured differently — local forums, Q&A communities, and media can carry more weight, and the competitive barrier often has not formed yet.

In practice: local-language GEO is not "translate the English page." It is constructing prompts in the target language, testing engine by engine (ChatGPT, Gemini, Perplexity, and AI Overviews have to be read separately), and only then deciding where content and external mentions should be built. Monitoring that runs English prompts only is a blind spot for a cross-border business.

Do not turn GEO into AI slop

GEO will become a buzzword, and it will be abused.

The most dangerous patterns:

  • Writing every article as a template Q&A of "how the AI would answer this."
  • Mass-generating vague paragraphs to blanket the long tail.
  • Dressing up unevidenced judgments in authoritative vocabulary.
  • Turning a personal anecdote into a repeatable formula.
  • Chasing mentions without caring whether the mention is accurate.

This is not only a matter of taste. Google's guide states it plainly: generating pages at scale to cover every query variation, where the main purpose is manipulating rankings or generative AI responses, can trigger the scaled content abuse spam policy. The correct answer to fan-out is one strong page covering several related intents, not a hundred thin pages each betting on one phrasing.

Content like that looks like GEO in the short run and damages trust in the long run. Answer engines do not need more identical answers; they need better sources.

If a piece carries no facts, experience, method, data, evidence, or clear judgment — if it just rewrites pages that already exist — it is not GEO. It is low-quality content in a new shell.

My conclusion

GEO is not SEO for ChatGPT.

More precisely, GEO is a re-examination of content quality in the answer engine era: is your content clear, verifiable, extractable, citable, paraphrasable, and explicit about its limits?

Traditional SEO makes a page easier to discover. GEO makes the information inside it easier to use correctly. The data from the first half of 2026 makes that concrete: source selection is no longer tied to the Top 10 (fan-out), being cited is not being seen (ghost citations), hidden markup loses to visible body copy (the schema study), and language markets are independent of each other (citation distribution reshapes by language). The direction is consistent — writing content as a verifiable node of information gets you closer to the point than researching any "AI preference."

For a personal blog that is good news. Content with long-term value should never have been chasing rank and clicks alone; it should be accumulating clear questions, real experience, stable evidence, and reusable judgment.

An article that does those things will suit an answer engine even if it was never written for one.

Update log

  • 2026-07-02: two corrections — (1) llms.txt downgraded from "you gain by having it, you lose without it" to "no effect on Google, a low-cost option elsewhere" (per the mythbusting list in Google's official guide); (2) "structured data clearly raises the probability of accurate citation" corrected to "no clear evidence that it lifts AI citations" (per Ahrefs' 1,885-page controlled study). Also added the query fan-out mechanism, the 38% Top 10 overlap, the 61.7% ghost citation share, the language variable, and 2026 sourcing throughout, plus three new sections: "Conclusions first," "Cited is not the same as seen," and "Language is an independent variable."
  • 2026-06-24: first published.

About the author

I am XCG (wowayou), working on English website content and Google SEO: first as an in-house English content operator for a software site, then in an agency SEO delivery group supporting 50+ B2B export sites, where I contributed to planning-stage work on about 20 projects plus existing-site checks and monthly GSC / GA4 reporting. This blog is my testbed — the robots.txt, structured data, and llms.txt configurations described above all run on this site. If you disagree or spot an error, the comments are below.

Comments →

CC BY-NC-SA 4.0

Comments

Comments are powered by GitHub Discussions. Sign in with GitHub to comment. Open the matching Discussion