Skip to content
Nikar

עבריתEnglish

How ChatGPT and Gemini decide which businesses to mention in an answer

· The Nikar team

Answer engines do not "rank" businesses the way a search engine does. They generate an answer from two sources: what was learned during training, and what is retrieved at the moment the question is asked, from a live search. A business enters an answer when it appears often enough, in a clear enough context, in the sources the model leans on — and above all when there is text tying its name to the question that was asked. That is why two businesses of the same size can get opposite results: size is not what decides, but how easy it is for the model to tie you to the need.

Two sources of information: training versus live retrieval

This distinction explains almost every behaviour that looks strange.

What was learned in training is what the model saw at the time it was trained. That explains why a model can talk confidently about a large, long-established company without searching for anything — the text about it was everywhere. It also explains why it sometimes talks confidently about a state of affairs that is no longer true.

What is retrieved at the moment the question is asked is the result of a search that runs at that instant. The professional term for it is Retrieval-Augmented Generation, or RAG: the system searches for relevant documents, passes them to the model along with the question, and the model writes an answer on the basis of what it received. That is what makes it possible for a business founded two months ago to appear at all.

These two paths have one practical consequence: they change at completely different rates. What is retrieved live can change within weeks; what was learned in training is updated only when a new model comes out, and nobody outside knows when. Anyone who names a date on which "the model will know you" is making it up.

If you arrived here without the background, what AEO is and how it differs from SEO.

What makes a model tie a business to a question?

This is where care is needed. The providers — OpenAI, Google, Anthropic — do not publish the exact rules, and the rules change. What can be said with confidence is what follows from the way systems like these work in general, and what we actually measure.

By what we measure, three things recur:

  • Co-occurrence. Your name appears in the same text as the words that describe the need — "דיני עבודה" ("employment law"), "חיפה" ("Haifa"), "פיטורים" ("dismissal") — and not only on your home page, where nothing but the name appears.
  • Entity clarity. It is clear that this is the same business everywhere. One name, not five. A business that appears in ten versions of its name is in effect split into ten weak businesses.
  • Corroboration from different sources. The same information appears in more than one place, and not only on the site you control completely.

Alongside these there is a technical layer: schema in JSON-LD format, following the schema.org standard, helps a machine understand what is on the page. Worth doing — but note the wording: it helps with understanding, and no provider has stated that it is a ranking factor in answers.

Example: two firms, the same quality, the opposite outcome

Take two law firms in Haifa, both good, both long-established.

The first wrote "ליטיגציה אזרחית ומסחרית" ("civil and commercial litigation") on its site, and appears in three places across the web under three versions of its name. It has not appeared in any article and in no ranking.

The second wrote a single page titled "פוטרתם בלי שימוע? מה החוק מחייב את המעסיק" ("Dismissed without a hearing? What the law requires of the employer"), appears in one version of its name everywhere, and is mentioned in one article about employment law.

When someone asks "מי עורך דין מומלץ לדיני עבודה בחיפה" ("who is a recommended employment lawyer in Haifa"), the second has three links between its name and the question: the words on the page, a consistent name that evidence can accumulate against, and an external source that corroborates. The first has none of them — not because it is any less good, but because no text ties it to what was asked.

That is the whole mechanism. It is not mysterious and it does not measure quality.

Why do the answers change from one run to the next?

Because models generate text through a process with randomness built into it, and on the live retrieval path the search itself can return different results at two points in time.

This is not a fault, and it is not something you can "switch off" from your side. It does say one very important thing: a single answer is not a measurement. Someone who opens ChatGPT once, gets their name in the answer and concludes that they "appear" has concluded from one sample. And someone who does not appear once and concludes that they have vanished — the same.

That is why we ask dozens of questions and not one, and repeat them over time. Not because it is more thorough, but because otherwise there is nothing to measure.

What about cited sources?

Sometimes the answer comes with links and sometimes it does not, and the difference says less than it seems.

When there are links, they are a good hint as to which sources were retrieved for that question — and that is useful information: if the engines cite a ranking or a guide you do not appear in, getting into it is usually the most worthwhile action there is.

When there are no links, that does not mean the answer was invented. It may simply be that it leaned on what was learned in training. The absence of a citation is an absence of information, not evidence.

It is also worth knowing that live retrieval does not necessarily go through Google. Some systems lean on other indexes — Bing is the most common of them — so a site that Google crawls well and another index barely knows can be reachable through one channel and not the other. That is another reason a good ranking in Google does not necessarily indicate visibility in the answers.

And one more point that is easy to miss: the links that appear in an answer are not necessarily everything that was retrieved. They are what the system chose to display. So a short list of sources is not evidence that few sources were read.

How does this work in Hebrew specifically?

Here there is a real difference, and it is the reason to read an Israeli article rather than a translation of an American one.

The volume of Hebrew text the models have seen is far smaller than the volume of English text. The practical result: every central Israeli source carries more weight. In English, a business that does not appear in one source will appear in twenty others. In Hebrew, the sources that cover a given industry in depth are sometimes five.

On top of that there is a transliteration problem. A business called "לוי ושות'" ("Levy & Co") can appear as "Levy & Co", as "Levi", and as a combination of the two. To a person it is obviously the same firm; to a machine it is three entities, unless something in the text ties them together.

There is a positive consequence too, and it gets talked about less: competition for a mention in Hebrew is far thinner than its English equivalent. In the American market, every niche is covered by dozens of detailed guides written precisely in order to get into the answers. In Hebrew, there are whole industries with not one page that answers directly the question customers ask. Whoever writes it first is not competing for a place on the list — they are creating the list.

And it is worth remembering that the questions themselves are asked in Hebrew. English content about an Israeli business helps less than it seems, because the link the model has to make is between your name and the words that appeared in the question — and those were in Hebrew.

It matters to say this: we measure through the providers' APIs, not through the app. A person talking to ChatGPT has a history, preferences and sometimes access to live browsing — so the answer they get can differ from the one we see. What the measurement does give is a consistent basis for comparison: the same questions, the same models, over time and against competitors.

What we do not know

This is the part most articles on the subject skip, which is why it is here.

  • We do not know the exact rules. The providers do not publish them. Anyone who declares that they know exactly how ChatGPT chooses sources is selling a certainty they do not have.
  • We do not know when a model is updated, and therefore we do not know how long it will take before a change you made affects the path that was learned in training.
  • We do not know how the two paths are weighted — when the model prefers what was retrieved over what it learned.
  • We do not measure every engine. We measure ChatGPT and Gemini, the two engines consumers in Israel actually use. Perplexity and Claude exist and are not in this check, for a reason of measurement validity and not of cost.

What can be done is to measure: to ask the same questions again and again, over time, and see what moves the needle. That is less impressive than a promise, and it is the only thing here that can be checked.

Want to see what the engines answer in your industry? The check is free and takes about a minute.

Frequently asked questions

How does ChatGPT choose which businesses to mention?
The answer is built from two sources: what the model learned during training, and what is retrieved at the moment the question is asked, from a live search. A business enters an answer when there is enough text tying its name to the need that was asked about, in the sources the model leans on. This is not a ranking by size or by quality of service — it is a question of how clear the link is.
Can I pay to appear in an answer?
No. As of today there is no advertising product that puts a business into a ChatGPT or Gemini answer. What makes a difference is where the business is mentioned across the web and what its site says.
Why does the answer change every time you ask?
Models generate text through a process with randomness built into it, and in some cases the live retrieval returns different results as well. That is why a single answer is noise and not a measurement — you have to ask the same question again and again over time in order to see a trend.
Does ChatGPT invent businesses?
Yes, it happens. A model that generates text can produce a name that sounds plausible and does not exist, or attach a detail that is not correct to a real business. That is one of the reasons to check what is said about you and not only whether anything is said at all.
What is RAG?
Retrieval-Augmented Generation. Instead of relying only on what was learned in training, the system searches for relevant documents at the moment the question is asked and passes them to the model along with the question. That is what allows a model to mention a business founded after training ended.
Does schema on the site help?
It helps a machine understand what is on the page, and that is almost always worth the effort. But no provider has stated that schema is a ranking factor in answers, and anyone who promises you that is selling you a certainty they do not have.