ChatGPT can find a page without citing it and talk about a brand without ever visiting its website: understanding how it reads information and picks its sources has become essential for anyone who wants to be visible in AI answers
As you read these lines, ChatGPT is recommending one of your competitors instead of you.
Hurts, doesn’t it? The good news is that in this new episode of SEO Confidential, which is an absolute bombshell, we have the person who knows exactly why.
Today we are joined by Suganthan Mohanadasan, one of the most authoritative and brilliant names in SEO and AI Search worldwide. Suganthan is the co-founder of the agency Snippet Digital and the “wizard” behind Keyword Insights, the content analysis tool used by giants like Amazon, IKEA, Adobe and eBay.
His work? One of a kind: he takes artificial intelligence apart piece by piece, runs experiment after experiment, uncovers the flaws and shares everything for free. A pure source of inspiration.
In this interview, Suganthan literally takes us behind the curtain of ChatGPT, Perplexity and the leading LLMs, gifting us an incredible amount of unpublished, remarkably precise and rather surprising data.
How do AI assistants decide what to read and what to cite? Why do they ignore your website?
If you want your brand to win over AI-powered answer engines, you need to know how these models reason behind the scenes, don’t you think?
So make yourself comfortable and get ready: what you are about to discover will make you look at things from a whole new perspective!

Why doesn’t AI cite your website? Unpublished data, mistakes and secret rules to win over the LLMs
Your study shows that ChatGPT can retrieve a page without citing it and can cite a brand without using its website as a source. What is the step that determines whether a page, once it has been found, is actually chosen as a source?
First of all, a caveat that applies to every answer I give here. All of this comes from analyzing the network traffic received by my accounts while I was logged in. The mechanisms are solid and can be reproduced in your own browser. The percentages, however, reflect what I observed on a single account, so they should be treated as indicative.
Three things can happen to a page, and they are distinct outcomes. Fetched means the engine retrieved the page and placed it in its working set of resources. Cited means a specific sentence in the answer links to the page as a source. Mentioned means the brand name appears in the text. You can win or lose each of these outcomes independently of the others.
In some measurements I have not published yet, ChatGPT fetched 3,554 pages across 57 conversations but cited only 110 of them. That is 3.1%. Almost everything gets read, but almost nothing earns a citation.
As for what separates a cited page from one that is merely fetched, the traffic shows three things. A citation is attached to a precise claim, not to a general topic. The page that wins, therefore, is the one that best supports a specific sentence in the answer, presenting the information early in the content, in plain HTML and including the relevant numbers directly.
Results are then deduplicated by domain: 20 pages from the same site covering the same topic can collapse into a single available slot, which means pages from the same site end up competing against each other.
The type of claim also shapes how the search is run. In one of my measurements I watched ChatGPT add the word “official” to 17 searches on its own. The system tends to look for your website when it needs to verify information and figures that concern you directly, while it turns to third-party sources when it has to establish, for instance, which product, service or solution is best.
Mentions, on the other hand, do not necessarily require the brand’s website to be consulted. In the same batch of measurements, 86 recommendations appeared without the brand’s website ever being fetched during the conversation. The model already knew those brands from the data used during training.
The final choice, that is, the reason one particular page is preferred over another, never reaches the browser. That process stays on OpenAI’s servers. As a result, we do not know the exact mechanism behind the final selection.
You discovered that for some questions ChatGPT does not search the web at all and answers straight from the model’s knowledge. How can we tell which queries trigger a web search and which ones leave practically no chance of earning a new citation?
The decision is visible in the traffic. Before ChatGPT runs a search, it classifies the question into a category, and one of these, internally labeled “text”, generates the answer directly from the training data, without performing any search.
In my June 2026 measurement, 3 of the 10 questions deliberately worded around current topics ended up in this category, including a request about the latest treatment guidelines for type 2 diabetes. You would naturally assume ChatGPT runs a search to answer a question like that. In that case, it did not.
It is the wording of the question that determines the category, not the topic. “Best 4K TVs to buy” triggered the dedicated shopping pipeline, while “best 4K TVs with reviews” kicked off a regular search. That is why I do not trust any rule of thumb for predicting what will happen, including my own.
How-to requests, definitions, code and translations tend to get answers based on the training data. Questions that call for an up-to-date figure, on the other hand, usually trigger a search. The diabetes example shows, however, just how relative that “usually” is.
The only method I feel comfortable recommending is testing your own queries directly. Open DevTools, ask the question and check whether any resource gets fetched. Repeat the test five times, because the routing is not deterministic. My free extension FanoutFox lets you view the same data with a single click.
When a query never triggers a search, no web page can enter the answer at that moment. The only way to appear in those answers is for the model to already have a memory of your brand through the training data. It is a slow process, built over years, including through archives such as Common Crawl.
What are the most common mistakes you see today in websites trying to earn visibility and citations on ChatGPT and the other AI engines?
The biggest problem is doing things in the wrong order. Teams buy technical audits, implement structured data and optimize llms.txt files without first checking whether the model actually associates their brand with its category.
In some data I am still processing, ChatGPT included brand names in its very first search query, before fetching a single page, in 21 out of 27 conversations. If your brand is not on that list, your server never even gets contacted and no amount of on-page optimization can change that. The fix is making sure your brand gets written about, reviewed and discussed over the years.
The second mistake is important information hidden behind JavaScript or buried inside images. The next two questions tackle exactly that, so I will cover it there.
The third mistake is publishing thin pages for every possible AI-facing query. Results are deduplicated by domain. Looking at the traffic you can see, for example, 19 pages from the same site grouped under the single page that won. One authoritative, comprehensive page for each specific claim beats 20 weak ones.
The fourth mistake is accidentally excluding yourself. Firewall rules and default CDN configurations that challenge every unknown bot can also block the commercial fetchers ChatGPT uses to retrieve pages, along with Common Crawl’s CCBot crawler. Until you run the proper checks, a block introduced automatically by a stock configuration can look like a deliberate choice.
There is also a general habit to avoid: treating a screenshot of these systems as a permanent truth. OpenAI changed the way sources appear in the traffic twice while I was conducting this research. That is why it is important to attach a date to every claim you rely on, including mine.
I built dedicated trackers precisely to monitor these changes:
How much does the way key information such as prices, features and product data is technically accessible on a website matter when it comes to being cited by ChatGPT? What happens when that information is hard for the AI to read?
I can answer using the model’s own words. In the June measurement, ChatGPT’s reasoning model described its own source-selection process while comparing a few SEO tools.
In the case of Ahrefs, whose prices are available in plain HTML, it consulted the official page, noting that the pricing page “seems more up to date, so I should cite that one”. For two other vendors, whose prices load via JavaScript, it observed instead that “the prices don’t appear directly in the search results, probably because they load via JavaScript”, adding that it could “cite third-party sources, since the official page is hard to parse”. In the end, those vendors’ prices were cited from G2 (a B2B software and services review and comparison platform, Ed.).
The entire consequence is visible in that single trace. The model wants to use the official page for information that concerns the company directly. It runs searches with the
site:operator on the pricing page and looks for currency symbols inside the fetched HTML.When parsing the page fails, though, it does not keep trying to retrieve that data. It relies instead on whoever else has published those figures, regardless of whether they are correct or up to date.
In essence, the accessibility of your pricing page determines who gets to report your company’s official information.
If prices, features or other important information are loaded via JavaScript, how hard does it become for ChatGPT to read them and use them as a source? What is the best way to make this data accessible to AI?
Two separate systems share the same problem. At answer-generation time, in my measurements prices loaded via JavaScript were not parsed correctly. The model itself said so in the reasoning I just quoted and, within a few seconds, it switched to a third-party source.
I cannot observe OpenAI’s fetchers from the inside, so I cannot claim they never execute JavaScript. What I can say is that I directly observed, on real vendors, both the parsing failure and the subsequent switch to an external source.
During the training phase, however, there is no ambiguity. CCBot, Common Crawl’s crawler, does not execute JavaScript, and Common Crawl is one of the foundational layers for much of the training data drawn from the open web. A page whose content only appears after hydration therefore gets archived as an essentially empty shell.
My checking tool retrieves the archived copy and compares it with the page currently live, letting you see which parts of your site CCBot has actually preserved.
The solution is rather old-school: prices, specs and key information should be server-side rendered as plain HTML text, preferably near the top of the page. It is best to avoid tabs and buttons that require a click to reveal a figure, as well as prices published only inside images or PDFs.
Then you need to check the page the same way machines see it: load it with JavaScript disabled and make sure prices, numbers and essential information are actually present in the HTML response.
If ChatGPT cannot find clear information on the official website and relies on external sources, how real is the risk that a brand narrative takes shape that differs from the actual one, potentially even damaging its reputation? And what can a brand do to prevent this kind of “parallel story”?
The risk is real because the switch to an alternative source happens automatically. These engines initially route factual claims toward official pages. ChatGPT adds the word “official” to its searches on its own and, in the late-July 2026 version, Perplexity attaches explicit trust hints to domains, limited to first-party information about the company’s products and services.
Both systems are therefore designed to let companies be the source of the information that concerns them. But when the page cannot be parsed correctly, or simply does not provide the figure being sought, the system turns to whoever does publish it. That is exactly why I watched the prices of two SaaS companies get cited from G2 rather than from their own official websites.
The worst case involves brands that publish no figures at all. If your enterprise pricing page just says “contact sales” while a forum thread speculates about a specific number, the model is looking at one source that contains a figure and another that does not. I have not directly documented a case like this yet, but all the mechanisms that could produce this outcome are visible in the traffic.
Prevention is unglamorous but concrete. Publish the information you want to be picked up, in plain HTML and on unmistakably official pages: prices, even as “starting from”, specs, dates and changelogs.
You also need to build a layer of external sources through reviews, comparison pages and Reddit, because opinions about your brand are sought elsewhere regardless of what you publish on your official website.
Then you need periodic checks. Every month, try the questions your potential customers ask and check which sources get cited for the claims about your brand. Answers change over time, and it is better to notice before your customers are the ones bringing that information to you.
ChatGPT can answer without searching the web, while Perplexity tends to run a search for every query. What does this difference mean for anyone trying to earn citations on the two platforms?
This splits AI visibility into two different games.
In my measurements, Perplexity never skipped the web search. The
skip_searchparameter came backfalsefor every query, both on June 25 and on July 21. The same question about how to change a flat tire, which ChatGPT had answered from its own memory, made Perplexity switch on its instructional mode, complete with a video card.This means that on Perplexity every query is potentially a visibility opportunity. Some resource always gets fetched and, as a result, there is always a slot that a web page or a video can occupy.
One important caveat, though: Perplexity’s default search uses a pre-crawled index rather than fetching each page directly from the site every time. Being present in that index, and having reasonably fresh content in it, is therefore the price of admission.
With ChatGPT, by contrast, a share of queries is closed to the web. No resource gets fetched, and so nothing you publish this quarter can directly enter those answers. In those cases, your presence depends entirely on what the model learned about your brand during training.
Query fan-out widens the gap further. ChatGPT rephrases the question far more “aggressively”: with its reasoning model it can generate 15 to 40 sub-queries, including brand names the system chooses on its own. Perplexity, in my measurements, searched for the literal wording of the question along with a couple of more conservative variants.
As a result, Perplexity tends to reward matching the words users actually type, while ChatGPT places more weight on the brand already being known to the model before the conversation even begins.
Informational, instructional or definition-focused content can therefore earn citations on Perplexity, while the exact same query on ChatGPT might not fetch a single web page.
One of the most striking differences involves YouTube: Perplexity often cites videos, while ChatGPT does so far less. What drives this difference, and what does it tell us about how the two platforms choose their sources?
It all comes down to what each pipeline is actually able to read, rather than a preference for a particular format.
A citation has to be tied to text the engine has fetched. When ChatGPT’s search hits a YouTube page, it gets the title and the description, but not the video transcript. As a result, it has no video text to cite.
In my June 2026 sample, YouTube was fetched 201 times and cited 0 times. Reddit, which is mostly text, was fetched 278 times and cited 11 times. It is a small sample, but Ahrefs observed a similar pattern across 1.4 million prompts: Reddit accounted for 1.93% of citations, versus 0.51% for YouTube.
Perplexity, on the other hand, processes video content. In my earbuds query, YouTube was cited 38 times, while in the flat-tire question three videos collected 22 citations between them. This happens through a dedicated video-answers card, which activates for practical how-to queries and product-related ones.
For brands, the takeaway is simple: the Reddit playbook many people learned by watching ChatGPT cannot be transferred automatically to Perplexity. For tutorial and product searches on Perplexity, a good video can play, in citation terms, the same role a text page plays on other systems.
Common Crawl is one of the great web archives used to build datasets for training AI models. How much can being present, barely present or completely absent from this archive affect a brand’s visibility, and why?
Common Crawl is the starting point for much of the data used to train models. When Mozilla analyzed the LLMs released between 2019 and 2023, it found that 64% had been trained, in some form, on data from Common Crawl. In GPT-3’s case, roughly 60% of the training data, by weight, came from filtered versions of Common Crawl.
Several of the major open training datasets, such as C4, RefinedWeb, RedPajama, Dolma and FineWeb, are also derived from filtered versions of this data.
Common Crawl also has a peculiar characteristic: the effects of its crawler’s visits can multiply over time. GPTBot fetches a page for OpenAI, while a page captured by CCBot can end up available to multiple parties training models on open-web data.
The link with AI visibility takes us back to my first answer. A model that puts your brand name in the query before it even runs a search must have learned that association somewhere, and this archive is probably one of the sources that helped build it. If Common Crawl has never captured your site, the model’s representation of your brand may have been built from pages published by others.
There are, however, two important limits readers should keep in mind.
Being in Common Crawl does not automatically mean being included in a model’s training data. Every lab applies its own filters to the collected data and, around 2023, the major model developers stopped publicly detailing the composition of their training datasets. So no one can guarantee that a given page in Common Crawl actually ends up in a model.
Common Crawl also samples the web partly based on how well connected a domain is to the rest of the network. Limited coverage, on its own, therefore proves little. My own site went from 2 captured pages to 55 in the space of a year.
The levers Common Crawl itself points to include links from well-connected sites, a sitemap and server-side rendering. The quality of how your site is represented in the archive matters too. The archived copy of my most-read article, for instance, was reached through an old link in a newsletter, tracking parameters still included.
In other words, the crawler follows the links that actually exist on the web, not the ideal version of the site we imagine we have built.
As you wrote on LinkedIn, roughly half a million sites block CCBot, Common Crawl’s crawler, and many may be doing so unknowingly because of CDNs, plugins or default settings. How much can unintentionally excluding yourself from one of the main data sources used to train AI models hurt a site?
Let’s start with the numbers. The HTTP Archive data shared by Chris Green in August 2026 shows that around 492,000 sites mention CCBot in their robots.txt file and that roughly 95% of those mentions block it.
Since July 1, 2025, Cloudflare blocks AI crawlers by default on every new domain it manages, while its managed robots.txt has reached 3.8 million domains whose owners never wrote one.
Squarespace uses its own default configuration, plugins ship blocklists, and a firewall can challenge CCBot even when the robots.txt shows no block at all, something no manual check can detect.
In 2026, then, a CCBot block is just as likely to come from a platform default as from a conscious decision. The BBC blocks it intentionally and appears 3 times in the July 2026 crawl. The problem is the sites that never made that choice.
The consequences come later and nothing flags them. Blocking CCBot today costs nothing: traffic does not change and no dashboard raises an alert. What you lose is presence in the monthly snapshots that new training datasets keep drawing from, and those datasets are kept up to date: FineWeb added six new crawls in the last year alone.
As a result, the models released two or three years from now learn from a web your site is not part of, and there is no button to resubmit content to a training corpus. I cannot quantify the damage as long as the labs keep their dataset compositions private, and neither can anyone else. What I can say is that the trade is one-way. Unless you are a publisher withholding content to gain leverage in licensing negotiations, the block gives you no advantage and the absence compounds over time.
Blocking CCBot does not change ChatGPT’s live answers, which retrieves information at question time through a different infrastructure. The damage is at the training layer, the one that determines whether a model already knows you exist.
My free checking tool queries Common Crawl’s archives directly, identifies the month a block first appeared and shows which template introduced it. When it detects a block, the date and the template usually make it possible to work out who made that decision. Often the answer is: nobody.
Make yourself understood by AI or you’re out!
Let’s wrap up. If there is one thing this conversation with Suganthan has made clear, it is that AI Search is neither magic nor luck: it is pure structure.
My take? We often get lost in overly complicated strategies or audits worth thousands of euros, only for ChatGPT to discard our page simply because it cannot read a price hidden behind some JavaScript code.
Ridiculous, isn’t it?
If your goal is to get cited and recommended by ChatGPT, Perplexity and friends, there is only one recipe: make life incredibly easy for the machines.
Put your key data in clean, plain HTML, make sure your brand actually gets talked about around the web (on Reddit, blogs and reviews) and check right away that you haven’t blocked the AI crawlers by mistake!
The truth is that the rules of search have changed forever. Those who understand today how these “big brains” reason will grab a huge slice of the market, while the ones sleeping soundly risk becoming invisible.
I hope this episode has opened your eyes and given you the right push to start seeing things from a new perspective. See you next week, right here on SEO Confidential, with another interview that is nothing short of mind-blowing.
#avantitutta
