SERRA – Strategia e lancio digitale

SEO Confidential: our exclusive interview with Olivier de Segonzac on how to win over LLMs

Written bySEO consultant & founder of SERRA

The most valuable resource for AI Search is already in-house: reviews, surveys, emails, chats and call center calls. That’s where you find your customers’ authentic language, their doubts and their needs

Press play and listen to what the interview with Olivier de Segonzac is all about

Have you ever wondered why ChatGPT keeps recommending the same brands, as if it had a pre-written list of trusted names? It’s no accident, and in this interview you’ll find out exactly how it works.

Welcome to a new episode of SEO Confidential, the series where we meet the best SEO and GEO experts on the international scene to understand what is really changing in the way AIs choose the sources they cite.

Today’s episode is one you can’t miss. We talk with Olivier de Segonzac, founder of the SEO agency RESONEO, where he also co-leads the Web Analytics and Data/Tech divisions.

If your goal is to become the source that ChatGPT, Gemini or Perplexity choose to cite when a user looks for information about your industry, this interview is a concrete roadmap to get there.

Olivier reveals how query fan-out really works, how much Wikidata matters even before Wikipedia, and why your content should answer your customers’ real-life situations, not your products’ technical specs.

The most useful takeaway concerns something every company already has in-house and almost nobody uses: customer reviews, call center transcripts, survey results, forum discussions, inbound emails…

That’s where you find the most valuable data in the AI Search era: the words people use to express their needs and describe their problems.

That is exactly where you need to start again to write content that an AI can find, understand and cite.

Olivier de Segonzac interviewed by SEO and GEO expert Roberto Serra
Olivier de Segonzac

The fastest route to getting cited by ChatGPT runs through dynamic visibility, while parametric visibility is built over time with Wikidata, digital PR and authoritative content

Hi Olivier, in your article on Search Engine Land, you write that ChatGPT now cites an ever smaller number of websites, focusing more and more on the sources it considers most authoritative. In practice, what allows a company to become one of those trusted sources the AI consistently chooses to cite?

Two things are essential: being reachable through ChatGPT’s retrieval chain, and being the page that gets opened when the system wants to verify a fact. On the first point, ChatGPT relies heavily on Google’s index, so the answer, however unexciting, is that ranking well on Google is still the price of admission. (This is changing, though. OpenAI is increasingly trying to reduce its dependence on Google and use a proprietary index, even if it cannot do so completely yet.)

On the second point, when ChatGPT digs into pages in depth, it behaves less like a search engine and more like a fact-checker. It gravitates toward primary sources such as official blogs, regulators, pricing pages and product pages that show current prices. When it looks for news about a company, it will often open the company’s own newsroom rather than press articles about it.

In practice, a trusted source is therefore the canonical, first-hand page packed with concrete data supporting a specific claim, presented in plain text. ChatGPT’s fetcher does not execute JavaScript, so anything rendered client-side is invisible to it. (The fetcher is the automated program that visits and reads web pages when the model decides to verify a piece of information or retrieve fresh data, Ed.)

Freshness matters too. We have observed that ChatGPT’s reranking system explicitly favors recent content, to the point that a recent but average-quality article can outrank an older, more reliable one. In practice, a page is considered trustworthy when it meets certain technical and informational requirements, not just because it belongs to a well-known site.

If a brand was not sufficiently represented in ChatGPT’s training data, can it still become a consistently cited source, or is Primary Bias a very hard obstacle to overcome? Which strategies can really make a difference?

Companies need to accept that they are operating on two different timelines at once. The parametric disadvantage is real: a model that does not know your brand will not naturally formulate search queries around it and will not favor your pages when ranking results. You start every conversation from a position of invisibility.

The dynamic layer offers a solution to this problem and remains fully accessible. A brand can rank on Google for the queries the model actually generates, secure mentions on the recurring set of domains cited in its industry, and build a presence on Reddit, which ChatGPT tends to treat more generously than many other sources. The set of domains regularly cited in a given industry is usually relatively small and can be measured.

The parametric gap can also close between one training cutoff and the next. What you publish today, including press coverage, a strong presence on Wikipedia and Wikidata, and pages captured by Common Crawl, could feed the next pre-training snapshot.

Dan Petrovic‘s work on brand recall in LLMs adds an important nuance. What determines whether a brand gets mentioned is not fame itself, but its centrality within the network of mentions. A mid-sized brand that is regularly cited alongside the right players in its niche can outperform a bigger but more isolated brand.

You distinguished between “parametric” visibility, meaning what the model already knows from training, and “dynamic” visibility, based on real-time web searches. Which of these two dimensions should a company invest in most to increase its chances of being recommended by ChatGPT?

Companies should start with dynamic visibility, because it can produce measurable results within a few weeks. Fresh, extractable content, indexing on Google, and mentions on the domains already cited in industry-specific answers can improve visibility on the observable layer, where the cited sources vary significantly from one month to the next.

Most companies still have obvious, low-cost improvements available here, starting with content that cannot be read without JavaScript.

Parametric visibility, however, determines whether a brand exists before any search is even performed. It affects which brands the model names spontaneously, which search queries it formulates, and which sources it prefers among the available results. Once encoded, this visibility remains relatively stable for months.

This cuts both ways. A strong position can be durable, but so can a weak position or complete absence. My recommendation is therefore to sequence the work rather than favoring one dimension over the other.

Companies should capture the immediate, dynamic wins right away, while continuing to work on parametric visibility through Wikidata, digital PR and authoritative English-language content. The benefits of that work may only show up after the next training cycle closes, and that timeline cannot be accelerated.

Many companies still optimize their sites mainly with Google in mind. Given how ChatGPT Search works, what should concretely change in an SEO strategy to increase the chances that a brand gets selected and cited in AI-generated answers?

Less than part of the GEO industry would have you believe, but more than most SEO teams are doing today. SEO fundamentals remain essential, because a page that is not indexed is excluded everywhere.

What changes is the unit of optimization. You no longer optimize just a page to earn a click. You optimize individual passages so they can be extracted and reused.

In practical terms, that means putting the complete answer in the opening paragraph of every section. It also means creating self-contained question-and-answer blocks, using precise statements with dates and numbers, adding comparison tables, and naming entities explicitly instead of relying on elegant synonyms. A paragraph might be quoted outside its original context, so it needs to remain understandable on its own.

Three specific points often surprise SEO teams. First, meta descriptions matter again. Large language models (LLMs) often process search results as a combination of title, snippet and URL, and we have observed that information added to meta descriptions resurfaces in AI-generated answers. The decisive information should therefore appear within the first 150 characters.

Second, all important content should be served in static HTML, since none of the major assistants reliably renders JavaScript. Third, strategic pages should also be published in English, since most of the internal search queries we detect are written in English, even when the user is French, for example.

There is more and more talk about GEO (Generative Engine Optimization). In your view, is it truly a new discipline destined to change how SEO is done, or is it mostly a new name for practices we largely already knew?

It is mostly a rebalancing of practices already known in SEO, combined with a few genuinely new elements. The retrieval layer is built on search engines, so indexing, rankings and snippets remain directly relevant. Structured content, entities and digital PR are not new either.

I am cautious about the rush to rename everything. GEO complements SEO. It does not replace it.

The genuinely new elements, however, are concrete. There is no traditional SEO equivalent of parametric visibility. SEO professionals never previously had to consider what an algorithm thought of a brand months after the information was published, with that representation encoded in the model’s weights.

Fan-out analysis is new too. The queries companies need to answer are machine-generated, and it is now possible to capture and study them. Measurement is another new discipline with its own pitfalls, since the system being measured can change overnight when the underlying model is swapped out.

GEO is therefore a new label for most of the work, but only a small part of it addresses phenomena that did not exist before. That small part is also the most interesting one.

Many business owners are starting to monitor their AI visibility using artificially constructed prompts. How reliable is this approach, considering that we still know very little about the real questions users ask AI systems?

The problem is not that the prompts are artificial. The problem is that most of them are guesswork. A consultant who invents 50 questions based on their own mental model of the client is mostly creating a benchmark of their own imagination.

The solution is to derive the questions from real customer data. Useful sources include call center transcripts, customer reviews, inbound emails, surveys and forum discussions. By collecting dozens of representative questions and varying them across different user profiles, companies can create synthetic prompts that credibly represent real usage.

Reliability therefore depends more on measurement discipline than on where the prompts come from. LLMs are not deterministic, so a single run is a sample, not a measurement. You need to run multiple executions and average the results.

There is also a structural fragility that many underestimate. Within a single product family such as GPT, successive models with different fine-tuning configurations and system prompts can behave very differently. The same prompt and the same brand can produce radically different KPIs from one day to the next, simply because OpenAI changed the default model.

We have observed a model change reduce the diversity of cited domains within a single day. AI visibility monitoring therefore needs to be continuous and measured separately for each model. A one-off audit is like a photograph of a moving train.

There is one final pitfall. ChatGPT internally uses many URLs that it never shows the user. Studies that count all of those URLs as citations produce flattering but inaccurate rankings.

Knowing that AI engines break a question down into multiple searches (the so-called query fan-out), does that change how a company should design its content?

A page that answers the main question is no longer the only goal. The goal is now the complete set of sub-questions the system derives from the original request.

In fast mode, ChatGPT may generate only a few queries. In reasoning mode, it can run dozens of queries across several search cycles. A brand’s visibility is therefore the sum of its presence across these intermediate queries, most of which the user never sees.

Two recurring patterns are relevant for content design. Most fan-out queries are written in English, regardless of the user’s language. Formats such as “best”, “top”, “vs” and “review” are also heavily overrepresented.

In practice, we analyze the recurring terms of fan-out searches within an industry and work them into titles, H1s, meta descriptions and opening paragraphs. Pages are then structured so that each sub-question gets its own self-contained, extractable block, rather than being buried inside a long paragraph.

The encouraging part is that, for common commercial intents, fan-out queries repeat frequently. The query space to cover is finite and observable.

The limitation is the same as with measurement. Fan-out behavior varies significantly across models and modes, and it changes when the underlying model changes. What was observed last quarter therefore needs to be verified again.

Some brands seem more concerned with creating content than with building their identity. How much does it matter today to be correctly recognized by Knowledge Graphs, and what concrete advantages can a brand gain in AI-generated answers?

It matters more than most roadmaps assume. ChatGPT performs Named Entity Recognition, and a brand either gets correctly identified as an entity or it does not.

Over the past year we have observed a steady expansion of its entity system. It now supports more entity types, better disambiguation, and entity cards displayed directly inside answers.

A brand whose products, technologies, services and executives are not structured and disambiguated on the web does not properly exist within that semantic layer, no matter how much content it publishes. Producing content without working on identity leads to mentions being misattributed or confused with similarly named entities.

The benefits are visible directly in ChatGPT’s interface. A recognized entity can receive its own card or sidebar, including images and structured information.

For local businesses, the connection is especially direct. ChatGPT’s location data is closely tied to Google’s, so a well-maintained Google Business Profile can show up on ChatGPT very quickly.

My working rule is simple: entities are your landmarks. Give your technologies, products and services a name, dedicate a specific page to each one, and state relationships and hierarchies explicitly, so that machines do not have to infer them.

For a company that wants to become a trusted source for ChatGPT, Gemini and the other AI systems, how important is it to have a complete and consistent presence on Wikidata? And what are the most frequent mistakes you see in how this information is managed?

Wikidata is very important and it is the starting point I recommend before Wikipedia. Wikipedia is among the most cited sources in ChatGPT’s answers, and without an entry a brand may be absent from a significant share of generated answers.

Editing Wikipedia directly is closely monitored and can be risky for brands. Wikidata is less restrictive and feeds many of the same knowledge graph systems. It helps distinguish, for example, “Apple, the company” from “apple, the fruit”.

There is another advantage many companies overlook. Google’s AI Mode relies on Knowledge Graph entities far more than traditional search results do. The same Wikidata work therefore also supports visibility on Google’s AI interfaces.

A solid minimum entry should include the founding date, headquarters, industry, executives, external identifiers and multilingual aliases. The company’s structured data should also contain a “sameAs” link pointing to the Wikidata entity.

The most common mistake is creating the entry and then abandoning it. That can lead to a page still listing the name of a CEO who left the company two years earlier.

Another recurring issue is inconsistent information across the website, LinkedIn, business registries and review platforms. Discrepancies in the company name, address or legal designation can undermine entity resolution.

Companies also often omit external identifiers and non-English aliases, even though these elements are essential for disambiguation. Another mistake is attempting direct Wikipedia edits that later get reverted and flagged, when the safer approach is usually to build reliable secondary sources first.

All of these mistakes stem from the same underlying problem: treating Wikidata as a one-off task instead of infrastructure to maintain.

Conversational searches often start from very specific needs and concrete situations. What should a company observe in its own customers to understand which content to create and which information to make available to AIs?

Customers describe situations, not technical specs. Nobody asks an AI for “a stroller with a 49-centimeter frame”. They ask for “a stroller that fits through subway turnstiles”.

The gap between how companies describe their products and how customers describe their problems is exactly where AI retrieval steps in.

The material needed to close that gap already exists inside most companies. It sits in call center logs, customer reviews, inbound emails, surveys and forum discussions about the product category. This corpus reveals real intents, actual vocabulary and the profiles of the people behind the questions.

Companies should therefore use this information in two ways. The first is content creation. Product and category pages should be rewritten around usage contexts: who the product is for, what it is used for, and in which situation it is relevant.

Each identified question should get its own self-contained answer block, with the answer stated in the first sentence. Product feeds are evolving in this direction too.

OpenAI’s shopping feed specifications include a Q&A field and take review data into account when ranking products. Customer language and customer feedback are therefore becoming ranking factors.

The second use is measurement. Those same real customer questions should become the prompts used to monitor AI visibility.

Companies that do this well stop asking: “What content should we produce?”

They ask instead: “Which customer situations have we not yet spoken to?”

Their own data can answer that question.

Your gold mine for AI Search? It’s already in your office

Did you see how many insights Olivier gave us in this interview?

He laid things out so clearly that by the end you can’t help thinking: so I was sitting on a pile of incredibly valuable information and didn’t even know it!

Think about it for a second. That email a customer wrote you to complain about something (or simply to clear up a doubt). That call center conversation where someone explained, in their own words, what they were really looking for. The survey you ran a year ago, full of your customers’ opinions, that nobody ever opened again.

All stuff you already have, just sitting there, and it’s worth more than any keyword research done at a desk. Because that’s where your customers speak plainly and express themselves in the most natural, spontaneous way possible.

And that is exactly the language an AI wants in order to find its way among a thousand brands and companies.

So the first step, really, is not to write more. It’s to go back and reread what you already have, and rewrite your content starting from there. Then sure, there’s more to do for SEO and GEO: Wikidata, structured data, Knowledge Graph, entities, static HTML — but start with what you have at hand.

A heartfelt thank you to Olivier, for his time and for the clarity with which he answered our questions. Insights like these, let me tell you, are priceless — so put them to good use!

See you at the next episode of SEO Confidential, with another internationally renowned guest. See you very soon!

The author

Roberto Serra

SEO consultant & founder of SERRA

Get the best updates on SEO

Already +6,200 professionals subscribed · one email a week, zero spam
✦ Let us start

Market, demand and competitors. Discover what your business could do. With the data in hand.

This is where your digital launch begins.