Skip to content
Praion
AEO · AI ENGINES

How ChatGPT search works — and what it means for you

Praion team · 12 min read

Published

QUICK ANSWER

ChatGPT can answer from the model's built-in knowledge or use web search. When it searches, OpenAI says it typically rewrites the request into one or more targeted queries and may run additional searches. To make a public site eligible for search results, allow OAI-SearchBot and requests from OpenAI's published searchbot IP ranges. No technical setting or content format guarantees a citation.

You ask ChatGPT "which accounting firm handles startups in Thessaloniki," and you get a clear answer, with two or three names and a few sources underneath.

The question we care about is simple: how did it pick those sources, and why isn't your own business among them?

ChatGPT doesn't always answer only from what it already knows, and it doesn't simply run a classic Google search. It can use web search, rewrite the request into more targeted queries, and synthesise an answer with citations. OpenAI doesn't publish the exact process used to select the final sources, while third-party research suggests that many pages can be retrieved without ultimately being cited.

Once you understand that process, it becomes clearer which parts you can control so your site is technically eligible and its content is useful and clear, without assuming that this will guarantee more citations.

Let's look at it in practice.

How does ChatGPT find its sources?

When a question can be answered better with up-to-date or specialised information, ChatGPT may automatically search the web. The user can also choose Search manually.

OpenAI says ChatGPT search typically rewrites the original request into one or more targeted queries sent to search providers. After reviewing the initial results, it may run additional, more specific searches.

In other words, it doesn't always search only for the phrase you typed. It explores different sides of the same question and uses, in the final answer, only a part of the sources it found.

Take the accounting-firm question as an example. The model doesn't just search for the phrase the way you'd type it into Google. It can break it into different sub-queries:

"accounting firms Thessaloniki," "accountant for startups," "cost of accounting support for new businesses."

Through these searches it identifies candidate sources and picks which information to use to compose the answer. OpenAI doesn't publicly disclose the exact process for judging and selecting sources.

What's most surprising is the large share of pages that end up unused.

In a third-party AirOps study of a dataset containing 548,534 pages retrieved while generating answers, about 15% ended up appearing as cited sources.

The other 85% had been surfaced during retrieval in that dataset but weren't cited in the final answer. This is useful third-party evidence, not an official or fixed OpenAI rule.

Having your page found isn't enough. The goal is for it to be judged relevant and useful enough to be picked as a source.

Where do these pages come from?

For web search, ChatGPT may work with third-party search providers. OpenAI doesn't present one provider as the fixed, exclusive source for every search. For a page's content to be included normally in ChatGPT Search summaries and snippets, OpenAI recommends not blocking OAI-SearchBot.

For the topic we're dealing with here, three user agents have different roles:

  • OAI-SearchBot is used for ChatGPT's search features.
  • GPTBot relates to the potential use of content for training OpenAI's generative AI foundation models.
  • ChatGPT-User may visit a page after a specific user action. It isn't used to crawl the web automatically or to determine whether a site can appear in ChatGPT Search.

This distinction matters, because each user agent has a different role.

It can do either.

In a typical response, ChatGPT can rely on the model's built-in knowledge or use web search.

The first is the model's built-in knowledge, which comes from its training data and training process. This knowledge isn't continuously updated from the web.

When web search is used, ChatGPT can retrieve more current information, cite sources and use material from them to compose the answer.

The difference is fundamental.

When the model answers only from the general knowledge it gained during training, it may name a company, if it has come across enough information about it. Usually, though, it adds no link and names no specific page as a source.

When search is used, the answer may include citations to pages used as sources, while the Sources panel can show those sources and other related links.

That's why the question "how often is ChatGPT's knowledge refreshed?" has two different answers.

The built-in knowledge depends on the data and the training period of each model. It isn't continuously updated every time something new is published on the web.

Search lets ChatGPT retrieve more recent information at the time of the question. This, however, doesn't guarantee that every new or updated page has already been found, or that it will be used in the answer.

If your content depends on current prices, new figures or market changes, you can't wait for it to be absorbed into the model's built-in knowledge. Keep the relevant page public and up to date so it can be retrieved when search is used.

It differs in three main ways, and each one affects how you should write your content.

First, the user doesn't just type keywords. They phrase a whole question in natural language.

Second, instead of first seeing a classic list of organic results, the user mainly gets a synthesised answer.

Third, the answer may include inline citations, while the "Sources" panel shows the sources used and, in some cases, additional related links. The number of sources isn't fixed and depends on the query and the answer.

In practice, this difference has an important consequence.

On Google, appearing on the first page of results can give you visibility and visits on its own.

On ChatGPT, the most direct visibility for a specific page comes when it appears as a citation or related link. A business can still be mentioned without its own site being used as the source, for example through a third-party page. The experience doesn't work like a classic paginated list of organic results.

The goal, then, isn't simply a spot in a ranking. For direct visibility of your own site, what matters is whether one of its pages is used or shown as a source or related link.

The AirOps study also found that, in its dataset, high Domain Authority wasn't a prerequisite for earning a citation: about 74% of citations went to domains with Domain Authority below 80.

That doesn't prove Domain Authority is a ChatGPT ranking factor. It does show that smaller or less authoritative domains can appear among the sources, and that citation visibility isn't limited to the largest sites.

What makes a page worth citing?

OpenAI doesn't publish a full list of factors or a specific formula for how sources are selected.

OpenAI says ChatGPT Search ranks results using multiple factors intended to help users find relevant, reliable information, and that placement isn't guaranteed. It doesn't publish a complete list or the weight of each factor. Beyond that, clarity, accuracy, recency where it matters, and technical accessibility are sound content and implementation practices — not published ranking factors.

The first practical area you can improve is structure.

When retrieving information, a search system can draw on specific passages of a page rather than necessarily the whole text as a single unit.

A clear question, with a complete answer from the first sentence, creates a self-contained and usable piece of content. It's a sound writing practice, but OpenAI doesn't document this specific format as a ranking or source-selection factor in ChatGPT Search.

Look at the difference:

❌ "Response time depends on the nature of the request, and we make every effort to serve our clients as fast as possible."

✔ "We reply to every message within one business day. If it's an urgent tax matter and you send it before 2 p.m., we reply the same day."

The second answer can be used as-is. It's clear, specific and understandable even without the rest of the text.

The first contains no specific information that can easily be used as a source.

The second practical area is the reliability of the content.

OpenAI says ChatGPT Search is designed to help users find reliable, relevant information, but the exact way a source's reliability is assessed isn't publicly known.

A site with a clear identity, well-documented content, a consistent presence and mentions from independent sources offers useful credibility signals for the reader. We can't, however, present those elements as published ChatGPT Search ranking factors. This doesn't mean smaller or newer sites are excluded.

The third practical area is recency, but only when it genuinely matters.

On topics that change, such as prices, comparisons, products or market figures, recently updated content is worth more than an older text.

On topics that stay stable, the date matters less.

Refreshing a page without a real change offers nothing. What helps is keeping the information that actually changes over time up to date.

Can you check whether ChatGPT cites you?

Yes. A first, indicative check can be done for free.

Open ChatGPT and ask a question the way a real customer of yours would. Ask about the field, the service or the problem you solve, not directly about your company's name.

For example:

"Which accounting firm handles startups in Thessaloniki?"

and not:

"What do you know about my company?"

Then look at which sites appear in the answer's sources.

A single answer, however, isn't a reliable measure of a business's overall visibility. The wording of the question, the language, the general location, the account's relevant memories and the timing can all affect the search and the result.

For a more reliable check you need a fixed set of real prompts, repeated measurements, and separate tracking of whether the business is merely named or whether a specific page of it is used as a source.

There are more technical ways to check too.

In your server logs you can check whether OAI-SearchBot visits your site and which pages it requests. The crawler's presence in the logs confirms that a visit happened. It doesn't prove, though, that the page has already been included in the search systems, or that it will be picked as a source.

ChatGPT automatically adds the utm_source=chatgpt.com parameter to the referral URLs of search results. So traffic reaching the site from related ChatGPT links can be recorded separately in your analytics.

If you need more systematic monitoring, there are tools that automatically put your customers' questions to different AI engines and record which sites get cited each time.

These tools are useful when you want to track a trend or spot when a competitor starts appearing in questions where your own business isn't mentioned.

How do you start appearing?

It takes two steps, and their order matters.

First you make sure the site is technically accessible to OAI-SearchBot. Then you make sure the content is accurate, useful and clear.

If the site isn't eligible for normal inclusion in search, good content on its own isn't enough.

On the technical side, the basic thing you can check right away is access.

Allowing only GPTBot isn't enough for your content to appear in ChatGPT search answers. GPTBot concerns the possible use of content for training OpenAI's generative models.

To manage your site's presence in the search features, you have to configure access for OAI-SearchBot separately.

If your robots.txt blocks OAI-SearchBot, OpenAI says the site won't be shown normally in ChatGPT search answers, although a URL may still appear as a navigational link.

The first check is to allow OAI-SearchBot in robots.txt. You should also make sure the host or CDN allows traffic from OpenAI's published searchbot IP ranges and that firewall or bot-protection rules aren't blocking it.

As for content, what applies is everything that makes a page worth citing.

Answer your customers' real questions.

Put the substantive answer at the start of each section.

Make sure every answer is self-contained, so it makes sense even without the rest of the text.

Build, step by step, the reliability that makes a source stand out, and keep the information that changes up to date.

None of this is a trick. It's how you create a clear, useful and reliable source.

What not to expect

Don't expect it to work like Google, and don't expect an immediate result.

There are three points you should know, because overlooking them can cost time and money.

The first is that ChatGPT search doesn't work like a classic search engine.

It's not enough to "come first," because there is no single first position in the usual sense. The answer uses a limited set of sources, and keywords alone aren't enough; the page needs genuinely useful, relevant information.

The second is that schema, on its own, isn't enough to put you in the answers.

Schema describes content in a structured format, but OpenAI doesn't publish documentation showing that adding schema on its own increases a page's chances of being selected in ChatGPT Search. In any case, it can't replace the actual information visible on the page.

The third is that there's no immediate result.

Changes don't lead to instant or guaranteed appearance.

OpenAI notes that it can take around 24 hours for its systems to adjust after a change to robots.txt. It doesn't, however, publish a guaranteed time for revisiting a page, retrieving it, or picking it as a source.

If you keep one thing, keep the right order of actions.

Let OAI-SearchBot read your site.

Write clear, self-contained answers to the questions your customers actually ask.

Build, step by step, the reliability that makes a source stand out.

The safest foundation is to publish information that is clear, specific, accurate and easy to access. That helps the reader and reduces ambiguity when a search system processes the page, without guaranteeing that it will select the page as a source.

At Praion we don't treat ChatGPT search as a separate channel. We work on the content and structure of the site so that it reads clearly for people and machines alike.

Because clear content is useful to people first, while also leaving less ambiguity for the systems that process it.

AEO aims to improve a business's ability to be found, understood and potentially appear in relevant AI answers — it doesn't guarantee a citation, recommendation, ranking or result. Answer engines continuously change how they select sources.

FREQUENTLY ASKED

Frequently asked questions

When a question can be answered better with up-to-date information, ChatGPT may search the web automatically. OpenAI says ChatGPT search typically rewrites the request into one or more targeted queries and may run additional searches after reviewing the first results. OpenAI doesn't publish the exact process used to select the final cited sources.

Yes. ChatGPT search synthesises an answer and may support it with citations and a Sources panel rather than showing only a classic list of organic results. The number of sources isn't fixed, and OpenAI says search results are ranked using multiple factors intended to surface relevant, reliable information. Placement is not guaranteed.

OpenAI doesn't publish a specific formula or the weight of individual factors. It says ChatGPT search ranks results using multiple factors intended to surface relevant, reliable information, and that placement isn't guaranteed. Clarity, accuracy and recency where it matters are sound content practices, not published ChatGPT ranking factors.

Each model's built-in knowledge has a cutoff and isn't continuously updated from the web. When web search is used, ChatGPT can use newer information at the time of the question. That doesn't mean every new or updated page will already have been discovered or that it will be selected as a source.

Yes, as an indication. Ask real questions about your field and see which businesses and sites appear in the sources, but don't treat one answer as a complete measurement. You can also check your server logs for OAI-SearchBot visits and identify ChatGPT search referrals in analytics through utm_source=chatgpt.com.

Start with technical eligibility: allow OAI-SearchBot in robots.txt and make sure your host or CDN doesn't block OpenAI's published searchbot IP ranges. GPTBot serves a different purpose and doesn't control Search inclusion. Then make sure the site contains accurate, useful and clear information. Technical access makes a page eligible; it doesn't guarantee placement or a citation.

No. GPTBot and OAI-SearchBot are controlled independently. You can allow OAI-SearchBot so your site remains eligible for ChatGPT search results while blocking GPTBot if you don't want your content used for potential training of OpenAI's generative AI foundation models.

Let's talk.

Half an hour to see whether ChatGPT cites your business today when someone asks about your field, which of your pages can be used as sources, and which information stays hard to find. No commitment.