Get Found Online

GEO Explained: How to Get Cited by ChatGPT and AI Search

Someone asked an AI assistant for an expert in your field this month and it named three people. Generative engine optimization is the unglamorous work of being one of them, and it has less to do with tricks than with being quotable.

Ievgen Krasovytskyi
Ievgen Krasovytskyi
AI & Automation · James Cook Media
· · 8 min read

Someone asked an AI assistant for an expert in your field this week and it named three people. Generative engine optimization is the unglamorous work of being one of them, and it has less to do with tricks than with being quotable.

Somebody typed your exact specialism into ChatGPT and asked who to talk to. They got a tidy paragraph, three names and a couple of links. You were not in it, and nothing about the answer told you why. There is no ranking to check, no position to track, no report that says "you came eleventh." Just an answer that happened without you.

That opacity is the new problem. Search at least had a scoreboard. This gives you a confident paragraph and no explanation, which makes it easy to sell you a solution and hard to check whether it worked.

So here is the honest version. Generative engine optimization, usually shortened to GEO, is the practice of making your expertise easy for an AI assistant to retrieve, quote and attribute. Not easy to admire. Easy to quote. That distinction is most of the work, and almost nobody explains it.

What generative engine optimization actually means

Generative engine optimization is the practice of writing and structuring your expertise so an AI assistant can retrieve it, quote a specific passage of it, and credit you by name. It differs from traditional SEO in its unit of success: SEO wins a ranked link, GEO wins a sentence inside somebody else's answer.

Search used to hand you a list and let the reader choose. An assistant reads the list on the reader's behalf, picks a handful of sources, writes one answer, and links out to what it leaned on. Google describes its generative features as using retrieval-augmented generation, meaning the answer is grounded in pages it has already indexed rather than invented from memory, and that it shows clickable links to the pages supporting the response (Google Search Central).

That is the whole opportunity. The assistant is not judging you. It is looking for material it can safely stand behind, then attributing it. Your job is to be the most convenient trustworthy thing in reach.

This is a different job from the one covered in plain-English SEO for coaches and consultants. That article is about being findable at all: pages, keywords, getting indexed. This one starts a step later. Once you are findable, what makes a machine choose your words over somebody else's?

How an assistant decides whose words to use

An AI assistant answers in roughly three moves: it retrieves candidate pages, selects the passages that most directly answer the question, and composes an answer with citations to the sources it used. The selection step is why a page can rank well and still never get quoted. It was retrieved, then nothing in it was liftable.

Nobody outside these companies can see the selection logic, and anyone claiming to have reverse-engineered it is guessing with confidence. But the shape of the pipeline is public: retrieval, then selection, then composition with attribution.

What is unusual is that a peer-reviewed paper actually tested this. In GEO: Generative Engine Optimization (KDD 2024), Aggarwal and colleagues built a benchmark of user queries and measured which content edits changed how much of a generated answer came from a given source. They report visibility gains of up to 40% from content-side changes alone, with the effect varying a lot by subject area (arXiv).

The interesting part is which edits worked. Adding relevant quotations, adding concrete statistics, and citing your own sources produced the biggest gains. Keyword stuffing did not help. The paper also reports the largest gains for sources that were not already sitting at the top of the results, which is the first genuinely encouraging finding in this field for a one-person business.

Treat that as one study on specific systems at a specific time, not a law of nature. But notice that every winning edit is something a real expert does naturally and a content mill does badly.

Comparison of an unquotable page and a self-contained quotable passage in AI search
Both versions can rank. Only the right-hand one survives being pulled out of the page.

Passage-level citability is the whole game

Passage-level citability means any single passage of your page can be lifted out and still make sense. An assistant does not quote your website. It quotes one specific chunk of it. If understanding your best sentence requires the three paragraphs above it, that sentence cannot travel, and it will not be used.

Read the answer block at the top of this section again. It is fifty-two words, it names its own subject, it never says "as we saw earlier", and it would be true and complete if you pasted it into a message to a colleague with no context at all.

Every section here opens with one. That is not a formatting quirk, it is the technique this article describes, done in public so you can see its shape rather than take my word for it. If an assistant ever quotes this piece, I would expect it to quote one of those blocks, because they are the only parts built to survive removal.

What makes a passage liftable

A passage travels when it can answer the question alone. In practice that means four things:

  • It restates its own subject. "This approach" is dead weight. "Passage-level citability" is a handle.
  • It commits. A hedge with three qualifiers is unquotable, because no one wants to attribute a shrug.
  • It carries something checkable: a number, a named source, a specific mechanism, a real constraint.
  • It ends. Forty to seventy words, then a full stop. Not a paragraph that trails into the next thought.

The uncomfortable implication is that most expert writing fails on the first two. We write in arcs, with the payoff earned over a page, because that is how you teach a human in a room. A retrieval system never reads the arc. It samples.

Where to put your best answer

Answer the question in the section where it is asked, immediately, before the nuance. Then add the nuance. You lose nothing with a human reader, who has always preferred the answer first, and you gain the only thing a machine can act on. One self-contained answer per real client question will take you further than most retainers do in a quarter.

Statement graphic reading you cannot make a model cite you, only worth citing
Three properties you control, in a field where almost nothing else is controllable.

What Google says you can stop worrying about

Google states that appearing in its AI features needs no special optimization, no new machine-readable files, no llms.txt, no Markdown, and no particular schema markup. Pages must simply be indexed and eligible to be shown with a snippet. Most paid GEO checklists sell work that the platform's own documentation says is unnecessary.

This is where a lot of money is currently being wasted, so it is worth quoting the source rather than paraphrasing it.

Google's guidance on optimizing for generative AI features says plainly: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search." It says there is "no requirement to break your content into tiny pieces for AI to better understand it." It says "You don't need to write in a specific way just for generative AI search." And on the schema question that half the GEO industry is built on, it says structured data is not required for generative AI search and there is no special markup to add (Google Search Central). Its documentation on AI features repeats the point: "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary" (Google Search Central).

Two honest caveats. That is Google talking about Google, and it does not bind ChatGPT, Perplexity or anything launched next quarter. And "not required" is not the same as "useless", since structured data still earns rich results in ordinary search.

But notice the trade. A self-contained answer to a real question helps you on every platform and helps the human reading it. An llms.txt file helps you on no platform that has documented reading one. When a tactic only makes sense to a machine, be suspicious of it.

Being verifiable is the part most people skip

An assistant will not attribute a claim it cannot stand behind. Verifiability means the specifics that let a system connect your name to a real, checkable person: a named author with credentials, dated pages, first-hand experience described concretely, and links to primary sources. Anonymous, undated, unsourced expertise reads as unattributable, however good it is.

Google's guidance asks for content that is not a commodity, warning against material that recycles what others have already said or that a model could have produced itself (Google Search Central). That is a liberating instruction for a solo expert. The one thing no content operation can manufacture is the specific, dated, slightly awkward detail of work you actually did.

The mistake I see most often is not laziness. It is modesty. The expert who has run four hundred difficult conversations writes "communication is important in leadership" because the specific version feels like boasting. The specific version is the only one that can be cited.

So write what only you could have written. What you tried that failed. The number you actually saw. The client situation, anonymised but real. The year. That is not self-promotion, it is your story made findable, and it is the format both a reader and a retrieval system need.

One small next step

See how the method works before you buy anything bigger

The AI MasterClass walks through the same approach we use with clients, including how to turn what you already know into passages worth quoting, with the prompts we actually use. Ninety minutes, at your own pace.

Take the AI MasterClass About 90 minutes · Watch at your own pace · No subscription

We have trained more than 100,000 experts since 2018, and the thing that separates the ones who get found is almost never budget. It is whether they were willing to be specific in public. That is a decision, not a spend.

The three doors you can actually open or close

You control three concrete things: whether AI search crawlers can reach your site, whether your pages allow snippets, and whether your pages are indexed at all. Blocking OpenAI's OAI-SearchBot removes you from ChatGPT search answers. A nosnippet directive removes the text an AI feature would quote. These are settings, not tactics.

Most GEO advice is soft. This part is not, and it is worth ten minutes with whoever manages your site.

Crawler access is separate from training access. OpenAI documents distinct crawlers: OAI-SearchBot indexes content for ChatGPT's search feature, while GPTBot gathers material for training models. Sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, and blocking GPTBot alone does not affect that appearance (OpenAI). Being quoted but not trained on is an available position, one robots.txt file away.

Perplexity draws the same line. PerplexityBot exists to surface and link websites in its results and is explicitly not used to crawl content for foundation models (Perplexity).

Snippet controls quietly gate you too. Google's AI features documentation lists nosnippet, data-nosnippet, max-snippet and noindex as the ways to limit what it shows from your pages (Google Search Central). A cautious developer who added nosnippet years ago to stop scrapers has, in effect, opted you out of being quoted.

Check those three things once. Then go back to writing.

Checklist of what to confirm and what to skip in generative engine optimization
The right column is where most GEO invoices go. None of it is required by the platforms.

What nobody can honestly promise you

No one can guarantee an AI assistant will cite you. Citation depends on the query, the model, the moment, and systems that change without notice, and none of that is purchasable. What you can control is whether you are quotable, verifiable and specific enough to be worth citing. That is the entire honest offer.

There is no submission form and no ranking to buy. Two people asking the same question can get different answers, and a source cited all spring can vanish when a model updates. Any vendor promising a guaranteed AI citation is either misunderstanding the mechanism or counting on you to.

What is stable is the demand underneath. These systems constantly need material that is specific, attributable and safe to stand behind. That favours the person with real experience over the operation with real budget, which has not been true of search for a long time.

You are probably already the best in the room and invisible online. GEO does not ask you to be louder. It asks you to be liftable: to say the specific thing, in a passage that stands on its own, with your name and the year attached. Do that fifty times in a year and you own the only asset here that compounds.

Questions people ask

What does GEO stand for, in plain English?
GEO stands for generative engine optimization. It means making your expertise easy for an AI assistant to find, quote and credit. Traditional SEO aims at a ranked link that a person clicks. GEO aims at a sentence inside an answer the assistant writes, with your name attached. The practical difference is that you optimise passages rather than pages.
Is GEO different from SEO, or a replacement for it?
It sits on top of SEO, not beside it. Google states that pages must be indexed and eligible to appear with a snippet before they can show up in generative features at all, and that the same foundational practices apply (Google Search Central). If your site is not indexed, no amount of GEO work matters. Fix findability first, citability second.
Do I need an llms.txt file to get cited by AI?
Not for Google, which says explicitly that no new machine-readable files, AI text files or Markdown are needed to appear in Search (Google Search Central). It costs little to add, but treat it as an experiment rather than a requirement, and be wary of anyone charging for it.
How do I know if ChatGPT or Perplexity is citing me?
Manually, for now. Write down ten to fifteen questions a real client would actually type, ask them in each assistant, and record which sources get named. Repeat monthly. It is crude, but it is more honest than most dashboards, because you are reading the actual output rather than a proxy score.
How long does it take before AI search mentions me?
There is no reliable published answer, and anyone quoting you a number is inventing it. What is knowable is the sequence: your pages have to be crawled, indexed, then chosen. So the first thing to check is access, not timing. Confirm the search crawlers are allowed and snippets permitted, then judge the content by whether it deserves quoting.
Can I do this without becoming a marketer?
Yes, and that is the part worth hearing. The winning edits in the research were quotations, statistics and sources, which is just showing your work. There is no persona to adopt and nothing to hype. You write the specific true thing you already know, in a shape that can be lifted out.

Keep reading