Syntra Systems
Cases Services Products About Blog IT Caravan
+998 70 010 68 44 +7 999 900 22 12
RusEngUzb
A human figure made of glowing blue particles on black, a metaphor for an AI answer assembled from many sources
Websites and platforms

GEO: how to get your site into ChatGPT and Gemini answers

By Natali Shumilovskaya · · 9 min read · updated

GEO (generative engine optimisation) is the work of preparing a site so that its pages reach the answers of AI services: ChatGPT, Perplexity, Gemini, Copilot and Alisa in Yandex search. It has three parts: open access to those services' search crawlers, write pages so that a direct answer can be lifted out of them, and check regularly what the answers say about your company. There is no magic here — only access, completeness and claims that can be checked.

Key takeaways

  • GEO comes down to two verifiable levers: access for the AI services' search crawlers and content a direct answer can be lifted from.
  • The search crawler and the training crawler are different: blocking GPTBot with the same line as OAI-SearchBot costs you appearances in ChatGPT answers.
  • Google states plainly that Search needs no llms.txt, no chunking of text and no special AI markup.
  • In the KDD 2024 experiment, quotations, statistics and source citations raised a page's visibility in generative answers by up to 40 %.
  • Yandex Webmaster reports the share of queries that mention your site in Alisa's answers, the cheapest objective measurement for a Russian-speaking audience.

How this differs from ordinary search optimisation

The object differs: in search you compete for the position of a page, in a generative answer you compete to have the model take your page as a source and link to it. The foundation is the same for both — the page has to be reachable by a crawler and answer the question.

What you set upOrdinary searchAnswers from AI services
Crawler accessGooglebot, YandexBotSeparate rules for OpenAI, Anthropic and Perplexity crawlers
What the reader seesYour title and snippetThe model's retelling with a link to the source
What you control directlyIndexing, structure, markupAccess and content only
What most often breaks itA blocked section or a duplicate pageA page with no direct answer in the opening lines

A full comparison of the two jobs, and the answer to where to start, is in our article on whether you still need SEO in the age of AI. In short: foundations first, visibility in answers second, because the second rests on the first.

There is a practical difference too: this is not a separate channel with its own budget and its own team. The same people edit robots.txt, the same page is rewritten around a direct answer, the same article gets a table and a checklist. The only genuinely new activity is the regular check of what the answers say.

How AI services reach your site

Through robots.txt, and each service runs several crawlers. The rule that matters: the search crawler and the training crawler are different, and one line cannot block both correctly. Block the search crawler and the site stops appearing in answers, however good its content is.

CrawlerWhat it works forWhat granting access gives you
OAI-SearchBotSearch inside ChatGPTYour site can appear in ChatGPT answers (OpenAI documentation)
GPTBotCollecting data to train OpenAI modelsInclusion in training sets; no effect on appearing in search
ChatGPT-UserFetching a page at a user's requestThe action is started by a person, so robots.txt rules do not apply to it
Claude-SearchBotSearch quality in ClaudeVisibility and accuracy of the site in search results (Anthropic documentation)
ClaudeBotCollecting data to train Anthropic modelsInclusion in training sets
PerplexityBotThe Perplexity search indexYour site can appear in Perplexity results (Perplexity documentation)
Google-ExtendedTraining Gemini models and grounding in Google appsNo effect on inclusion in Google Search or on ranking (Google documentation)

It is worth checking separately whether bot protection on the hosting side or the CDN blocks these crawlers: a robots.txt rule may allow access while a network filter quietly returns an error to them.

Mistake A company blocks GPTBot so its texts stay out of training, and with the same line blocks OAI-SearchBot too. That will not stop training — the pages exist in other sources anyway — but the site will disappear from ChatGPT answers.

What actually improves the odds of being used in an answer

There are two verifiable levers: access and content. Everything else sold under the name “AI optimisation” is either unconfirmed by the platforms or contradicted by them outright.

Google has been unambiguous about it: to appear in Search you do not need to create machine-readable files such as llms.txt, chunk your text or chase inauthentic mentions; the generative features rest on the same ranking systems, and no special markup is required for them (Google Search Central guide).

On the content side there is a public experiment. In the paper “GEO: Generative Engine Optimization” (Aggarwal et al., KDD 2024) the authors tested a set of techniques against generative search engines: adding quotations, statistics and source citations raised a page's visibility in answers by up to 40 %, and the effect varied noticeably by subject area (the preprint).

From this follows a practical list — not platform rules, but what agrees both with their documentation and with the experiment.

It helps to separate two different outcomes. A mention is when the model names the company in the text of the answer; a citation is when your domain is linked next to it. The first builds recognition, the second brings people to the site, and they are earned differently: mentions through the presence of the brand across sources on the topic, citations through your page holding the answer that is needed.

What a quotable page looks like

Such a page answers the question immediately and then opens it up part by part. The routine is simple and repeats from article to article.

  1. Write the reader's question word for word. Not a topic but a question: “how much does it cost”, “X or Y”, “how do I connect it”.
  2. Answer in the first two or three sentences. The answer has to make sense without the rest of the text.
  3. Break the question into sub-questions. Each heading is a question of its own, and the first sentence under it is the answer.
  4. Back it with figures and links. Every figure needs its source in the same paragraph, otherwise it will not be picked up.
  5. Add a comparison or a sequence of steps. Tables, steps and checklists are easier to lift out than solid prose.
  6. Close the remaining questions with an FAQ. Short self-contained answers that do not repeat the headings.

A word on length. There is no rule that says “write three thousand words or more”: Google states plainly that Search has no preferred word count (Google guidance). Length follows from the completeness of the answer — extra paragraphs make a page worse, not better.

Schema.org markup: what still works and what does not

Markup helps machines read facts off the page unambiguously: who the author is, when it was updated, which organisation it belongs to, what the price is. For Google's generative features it is not required, and no special “AI markup” exists — the same Google guide says so.

Type names have to be written correctly or the markup simply is not read: a list of questions and answers is FAQPage, breadcrumbs are BreadcrumbList, a company is Organization, an article is Article, a person is Person.

FAQ markup no longer produces a rich result, though: in May 2026 Google announced that the FAQ rich result would stop appearing in Search (Google documentation changelog). Keeping FAQPage still makes sense — other systems read it — but not for the sake of how the listing looks.

Note Markup does not turn a claim into a fact: it only helps a machine read what is already on the page. If the page carries no author, no date and no specifics, markup will not create them.

Yandex and Alisa: what to account for with a Russian-speaking audience

In Yandex search the AI answers are produced by Alisa, and Yandex has described the source selection in its webmaster help: the answer is built on the most relevant and highest-quality pages, which usually already rank well in search (Yandex Webmaster help).

The same place now offers a measurement tool: Webmaster reports share of voice — the share of queries mentioning your site among all queries Alisa answered — along with sample queries, pages and competing sites. For an Uzbek business serving a Russian-speaking audience this is the cheapest way to get an objective figure.

Hence the practical argument for a full Uzbek version rather than an interface toggle: on many business topics there is little material in Uzbek on the open web, which gives a well-made page a better chance of ending up among the sources of an answer.

A technical checklist

Before turning to content, check that the site is reachable at all by the systems you are preparing it for.

The point about scripts is worth checking by hand: open the page source and look for the first paragraph in it. If you cannot find it, a crawler may not find it either.

How to measure the result

You measure presence in answers rather than a position: the share of answers mentioning the brand, the share that link to the site, the comparison with competitors and the accuracy of what is said about you. All of it has to be counted against a fixed set of control queries, otherwise the numbers cannot be compared month to month.

Instability is a separate difficulty: the same query returns different answers, so a single check proves nothing. The structure of a monthly measurement, the metrics and their formulas are covered in our article on the AI visibility report.

Some of the numbers come from the platforms themselves: Google Search Console reports impressions in the generative features of Search, Bing Webmaster Tools reports citations in Copilot answers, and Yandex Webmaster reports share of voice in Alisa's answers. That is not the full picture: ChatGPT, Perplexity and Claude give site owners no such reports, so for them the measurement stays manual.

How we do this

Syntra Systems does not sell “promotion inside neural networks” as a separate service with promised mentions: the platforms do not publish their ranking rules, so there is nothing to promise. Basic SEO and GEO setup is part of building an online store, from $10,000; the scope is described on the website development page. What our projects look like can be seen in the case studies.

After that it is long-term work: crawler access, the structure of articles, regular measurement and reworking of pages that answers ignore. It sits inside project development, from $1,500 a month, with a fixed number of team hours and a quarterly review of priorities (technical support).

Let’s discuss your project

Tell us what you need, and we will estimate the timeline and cost and suggest a solution.

Discuss website development

Frequently asked questions

Will GEO replace search optimisation?

No, it sits on top of it. A page reaches a generative answer only once it has been fetched, read and judged relevant, and all of that is the job of ordinary optimisation. Without the foundation there is nothing to build on.

Does a site need an llms.txt file?

For Google Search, no — its guide on generative features says so. Other services describe access control through robots.txt and their own crawlers rather than an extra file. Adding llms.txt does no harm, but do not count on it.

How do I check whether ChatGPT sees my site?

Ask it a dozen questions you expect to be found for and see whether the company appears in the answer and whether the domain is linked. Repeat the check several times: answers are unstable, and a single run proves nothing.

Should I block the crawlers that collect training data?

That is a decision about rights to your content, not about visibility. Blocking a training crawler does not affect appearing in that service's search, nor does it remove data already collected. The point is not to hit the search crawlers by accident: their robots.txt lines are written separately.

Why does an AI service say untrue things about my company?

The model assembles an answer from whatever sources it has, including outdated pages, directories and second-hand retellings. The fix is not a complaint but publishing unambiguous facts on your own site and updating the business listings where the data has gone stale.

When will AI models start mentioning us in their answers?

We won't promise a date: platforms don't publish their source-selection rules and decide for themselves when to recrawl a site. Progress shows up earlier in reports — Yandex Webmaster shows your share of voice in Alice's answers, Search Console shows impressions in Google's generative features, and Bing Webmaster Tools shows citations in Copilot.

Do we need a separate page for every question?

Google states directly that breaking text into fragments for generative search isn't necessary. A separate page makes sense when a question has its own audience and its own intent — say, "how much does it cost" versus "how to choose." Sub-questions on the same topic can stay as headings on a single page.

Read also