Google Discover, AI Overviews and Information Gain: John explains how publisher visibility is changing, the value of AI citations, and the strategies needed to earn the authority you deserve
Have you ever wondered how much it is really worth to appear inside an AI-generated answer?
It is a question every business owner should be asking today, because the answer is far from simple.
On one hand, there is brand visibility, which matters even without a click because it teaches AI to associate your name with your industry. On the other, there is the risk of celebrating a citation that sends no one to your website. Understanding the difference between these two things is precisely one of the themes explored in this interview.
Today on SEO Confidential, we are joined by John Shehata, one of the world’s leading experts in news SEO. He is the CEO and founder of NewzDash, the real-time analytics platform used by BBC, NBC News, CBS News, Reuters, Le Figaro and, in Italy, by GEDI and Quotidiano Nazionale.
He created GDdash, which measures brand visibility across AI platforms. He previously led audience development at Condé Nast across 60 brands in 11 markets and search at Disney ABC. Together with Barry Adams, he organizes NESS, the News & Editorial SEO Summit, taking place this year from October 13–15, 2026.
Let me warn you right away: this is one of the most in-depth interviews ever published on SEO Confidential, offering a rare level of insight into Google News and Google Discover, and much more.
John goes into remarkable detail and shares a number of best practices, explaining them with such clarity that they are accessible to anyone: which metrics you should start tracking immediately, how Google Discover really works, which signals actually matter for getting cited by AI and which ones you are simply chasing for no reason.
Get ready to take notes. Lots of them.

Ranking and citations? Then invest in Information Gain!
With the growth of AI Overviews, many news stories and pieces of content are now summarized directly on the search results page. In your view, what is the best way to measure the real impact of these systems on publisher traffic and visibility? Which metrics should publishers and SEO professionals be monitoring today?
The mistake most publishers make is trying to measure AI Overviews using the metrics they already use. Clicks and sessions tell you what happened afterward. They don’t tell you why it happened, and they don’t tell you what you’re missing on queries where you have never ranked.
Start with a control. Take the queries where an AI Overview appears and compare them with comparable queries where one does not, analyzing impressions, CTR, clicks and position over time. Even better, take the same query before and after it started showing an AIO. Without this comparison, you risk attributing to AI phenomena that could actually be caused by seasonality, a ranking change or a slow news week.
As of September 2026, AI Overviews appear for roughly half of the trending news queries we monitor on NewzDash, with rates reaching 80% in categories such as health. So the first question isn’t “has my traffic declined?” but “how exposed am I?”
Answering that requires a different set of metrics:
- AI Overview activation rate across your topics, broken down by section. Health and science behave completely differently from entertainment and sports, so an overall site-wide figure hides what is actually happening.
- Citation rate and citation position. It’s not enough to know that an AIO exists. You need to know whether you are cited, whether you are the first or sixth source, and which competitors are cited more often than you.
- Top Stories positioning relative to the AIO. This is the metric almost nobody measures. Our data shows that, among trending news queries that generate a Top Stories carousel, 16–17% in the US and UK now show that carousel inside the AI Overview rather than as a separate module: 15.5% in the US and 17.46% in the UK. When this happens, the headline, image and brand appear inside the dominant search feature rather than below it. When they don’t, a standalone carousel can be pushed much further down the page. With the same ranking, the result in terms of clicks can be very different. That’s why you need monitoring that tells you which of the two positions you achieved for each query, not simply that you ranked.
- CTR by query type in Search Console, compared over time, for queries where an AIO is present versus those where it isn’t. Google’s new generative AI performance reports can be useful here, but GSC data needs to be read alongside SERP monitoring, not instead of it.
- Where exposure is concentrated. Entertainment exceeds 30% for embedded Top Stories in both countries. Health and science are close to zero. If your business depends on a particular category, that category’s figure is the one that matters to you.
At that point, keep three elements separate: visibility, traffic and business outcomes. A citation in an AIO represents brand exposure and has a certain value, but an impression is not equivalent to a visit. If AI visibility increases by 40% and traffic falls by 20%, that does not automatically mean the result is positive or negative. The question is what happened to subscriptions, registrations and “returning readers”.
Analytics tools will always remain the source of truth for overall click data. What they cannot show is the query, the SERP configuration and the set of competitors present at the exact moment a user searched. That’s why we built NewzDash around lead indicators: rankings and visibility across Top Stories, organic results, Discover and AI Overviews, updated every few minutes, including information showing whether a Top Stories result appeared as a standalone module in the SERP or inside the AI Overview. This allows the SEO team to see exposure before it turns into a traffic loss in the monthly report.
If Google answers certain searches directly through AI, should publishers continue investing in those topics in the same way, or shift resources toward subjects where a click through to the site is still more likely?
Change, yes, but based on the type of query, not the topic.
AI wins on the simplest and most standardized information. A query such as “what time does the game start?” or “who won the election in region X?” can be answered by an AIO, while the hundreds of outlets that have all written the same 300 words end up sharing almost nothing. These were low-value pieces of content even before AI. The difference is that Google still distributed some of the clicks. Now it often doesn’t.
That doesn’t mean abandoning the topic. For two reasons. First, basic coverage of a topic allows Google and AI systems to understand which subjects your publication is authoritative on, and that authority can subsequently help you get cited and ranked across Google’s different surfaces. A topic you have never covered is a topic for which you can never be considered a reliable source.
The second reason is that the useful question isn’t “can AI answer this question?” but “what happens after the answer?” If the entire value of an article is contained in three sentences, that content is exposed. If the reader still needs the interview, the data, the local consequences or real-time updates, the AIO is a doorway, not a wall.
Publishers therefore need to be honest about which part of the topic they are actually covering. Breaking news continues to win Top Stories and the Live label during the first few hours. What no longer supports a newsroom is simply rewriting a wire story without adding anything. The follow-up story, the analysis, the expert quoted by name, the original data, the local perspective: these are the types of content that withstand the AIO and can be cited within it.
So this is the practical approach we recommend to publishers:
- Continue following the news during the first few hours, when speed can help you win Top Stories. Continue covering developing events through live updates as well.
- Don’t expect the first article to generate traffic for three days. Plan a second, more in-depth piece, and decide whether to update the same URL on the same day or link the two articles together.
- Analyze your data by section. Where the AIO activation rate is high and the citation rate is low, the rewriting strategy isn’t working. Where the AIO is present but your Top Stories results are embedded within it, that remains a strong position worth competing for.
- Allocate some of the resources freed up to other sections and to Google Discover. For many of the publishers we work with, Discover already represents 70% or more of their Google traffic and rewards a different type of content: strong entities, effective images and a clear reason for “why now”.
Each news section should then be analyzed separately, because exposure is not distributed evenly. Our NewzDash SERP data shows that 60–80% of trending health queries now feature an AI Overview. At these levels, the issue is no longer simply SEO.
It becomes an editorial and business question: is this section part of the brand’s identity? Does it generate a return that justifies the costs? And does that return necessarily have to come from search?
A health desk that feeds the newsletter, podcast and subscription journey may be worth maintaining even if Google sends it very little traffic. A health section created solely to generate search traffic, on the other hand, probably isn’t sustainable. Sports and entertainment sit at the opposite end of the same scale and, for each, the answer will be different.
No one can make this decision for the newsroom, but the newsroom cannot make it without having access to data at the individual-section level.
The issue of resources is really a portfolio issue. Most teams still fund content based on topic. They should instead fund content based on the surface on which it can realistically compete, section by section.
Is there a mistake you see publishers repeatedly making as they try to adapt to AI Overviews and AI search?
Treating “AI search” as a separate discipline, with a dedicated team, a specific checklist and a series of ad hoc tricks.
The fundamental principles that help you win space in Top Stories are largely the same ones that support a citation in an AIO: a clear headline containing the relevant entities, an inverted-pyramid introduction that states the fact within the first 100 words, well-structured pages, fast indexing and a brand that Google already trusts for that category.
Publishers chasing some mythical AI ranking factor while struggling with slow indexing and vague headlines are addressing the wrong problem. The worst version of this approach is writing for the model instead of the reader: adding summary blocks and “key takeaways” boxes to every article in the hope that an LLM will pick them up. The model doesn’t need help summarizing. It needs something worth summarizing.
The second part of this mistake concerns measurement. I see teams celebrating because “we appear on ChatGPT” without asking whether that query has significant volume, whether the AI has associated the brand with the topic they actually want to own, or whether that presence has generated any traffic. AI visibility is not a single metric. It means understanding whether AI platforms associate your brand with your topics, whether they cite you and whether that citation sends anyone to your site. These are three different levels and require three different measurements.
Within this framework, two distinctions matter. The first is separating branded prompts from non-branded ones. Appearing when someone asks “what does the BBC say about the strike?” proves very little. Appearing when someone asks “what is happening with the strike?” represents the kind of visibility that can help build an audience, and most publishers have never measured this second phenomenon.
The second distinction is between retrieved sources and cited sources. Platforms retrieve far more sources than they ultimately display. A brand that is retrieved but not cited is read by the model but loses attribution to someone else. That’s a very different problem from not being found at all and, in our experience, more often indicates a problem with content clarity and extractability than a lack of authority.
That’s the model we integrated into GDdash’s AI tracking, because a screenshot of a single prompt isn’t a strategy.
Google Discover has become a very important traffic source for many publishers. If Google starts placing more and more AI-generated content inside Discover as well, do you think this could reduce traffic to original articles? And what should publishers do to prepare for this possible change?
Yes, it can, and publishers should consider it the riskiest change, more so than AI Overviews.
Search is driven by intent. Discover is driven by interest. That difference makes a summary on Discover more damaging: in Search, the user had a question and the AIO provided the answer. On Discover, the user had a curiosity and the card represented the click. If the summary satisfies that curiosity, the click doesn’t happen.
The rest is simply mathematics. In our analysis of more than 400 publishers, Discover’s share of Google-generated mobile traffic almost doubled in two years, rising from 37% in 2023 to 67.5% at the end of 2025, while web Search fell from 51% to 27%. Google is already experimenting with AI-generated summaries and experiences that group news within Discover. A citation in an AIO that you didn’t get costs you a click on a query. A Discover card replaced by a summary can cost you your primary traffic channel.
There is also a control issue that most publishers have not yet fully understood. Google’s new Search Console control for generative AI allows eligible sites to opt out of AI Overviews, AI Mode and generative Discover features. There is no option, however, to block the AI summary while keeping your link and attribution within the same feature. Opting out means no link, no impression and no traffic from that experience, while other publishers continue to appear. The old workaround, the nosnippet tag, also creates what I call a semi-block on Discover because it removes the descriptive text the card needs to be displayed. Based on our observations, this reduces Discover reach much more than would be necessary to block only the AI features it was intended for.
How to prepare:
- Know your Discover exposure by category. Entertainment lost a significant share of its visibility at the end of 2025, and the February 2026 Discover core update redistributed winners and losers at the domain level. If you don’t monitor your position within the category against the top ten sites, you cannot distinguish an AI-related change from normal volatility.
- Become the source worth clicking even when a summary exists: original reporting, named entities, images at least 1,200 pixels wide and headlines that state the news instead of merely creating curiosity. Curiosity-gap headlines encourage the system to penalize you through quick exits and short sessions.
- Use data-nosnippet on the specific paragraphs you want to protect, rather than applying nosnippet across the entire site.
- Decide your opt-out policy based on numbers, not opinions. Measure how much traffic comes from generative Discover features before disabling anything.
- Diversify. Newsletters, apps, direct traffic and creator distribution provide protection against dependence on a single Google surface. Use Discover aggressively, but treat it as a distribution channel, not the business model.
Discover is distribution engineering. Publishers that build a strong pipeline will continue to earn the card, regardless of whether the summary is present.
How much can we really understand today about how Google Discover works by observing the signals Google collects from users’ devices? And which aspects of its algorithm remain impossible to know from the outside?
More than most people imagine, but less than the vendors selling “Discover secrets” would have you believe.
The most useful work on this subject has come from analysis of Discover SDK telemetry conducted by Metehan Yeşilyurt, which we expanded at NewzDash by turning it into a nine-step content pipeline: acquisition, Open Graph analysis, classification, collection-level filtering, interest matching, server-side ranking, feed composition, distribution and feedback loop.
The value of this perspective is not that it reveals a ranking secret. It is that Discover can fail at different points in the pipeline. A site-wide visibility drop is often caused by a rendering issue, such as an og:image or og:title not being correctly processed after a template change, or by an editor-level filtering issue rather than the headline. Most teams analyze Discover as if it were Search, and that’s the trap.
What we can observe with a good degree of reliability:
- Which entities, categories, card types and image variants actually appear and for how long. In our DiscoverPulse data, articles generally remain in Discover for one to three days.
- Whether Google has rewritten the description, changed the image or grouped the content together with that of a competitor.
- Domain visibility over time and which sites enter or leave the feed in a given market. Our DiscoverPulse panel is based on approximately six million consented users in the United States and one million in the United Kingdom, mapped to Google’s taxonomy, which includes around 2,000 categories and more than nine billion entities. It is a representative sample, not an exhaustive one, and I always recommend that publishers read it alongside their own Search Console data.
- The fact that telemetry indicates a ranking system based on a model that predicts click-through rate and runs on Google’s servers, fed by headline, image quality, content freshness and the historical relationship between clicks and views.
What remains invisible from the outside:
- The weight assigned to each factor for each individual user. Two people can see different feeds at the same time, so a single snapshot can be misleading. We also cannot see how Google balances a short-term interest, such as a single recipe search, against a long-term interest, such as a football team followed for years.
- The internal mechanisms that determine collection-level filtering. We can see the result, not what triggered the filter.
- The actual pCTR model (predicted CTR, Ed.), and how post-click behavior, such as short sessions and closing cards, negatively affects a domain.
- Impressions broken down by card type. Google provides Discover data in terms of clicks and impressions, but not by position or element type. Without panel data, it is therefore impossible to separate story quality from the effect of position in the feed.
- Volatility caused by experiments. Content distribution is an active system using streaming, caching and background synchronization, so some drops may be caused by temporary changes and do not necessarily represent a negative judgment.
The most accurate position is this: we can measure results well, we can clearly analyze the different stages of the pipeline, but we cannot directly confirm the weights used in ranking. Reverse engineering can help us understand how the machine appears to behave. It cannot give us access to its source code.
The useful question, though, has never been “do we know the algorithm?” It is “do we know enough to make better editorial decisions?” Today, the answer is yes.
Google Discover increasingly seems to rely on predicting users’ interests and behavior. In your view, how important are the signals collected from interactions with Search, YouTube, Chrome and other Google services in deciding which content to show a person?
They are the foundation of the system, and Google states this in its own documentation. Discover personalization is based on Web & App Activity, which includes search history, and YouTube History when that feature is enabled, as well as location, topics and brands a user follows, and content they interact with or explicitly reject in the feed.
Search shows what a person is actively looking for. YouTube reveals interests that never emerge from a query. Location changes which news is actually relevant. The reason Discover can show someone a story about a niche football team or a specific chip manufacturer is that Google already knows, through Search and YouTube, that this person is interested in that topic.
Publishers’ mistake is thinking they can act on these signals at the individual level. You cannot optimize for John or Roberto. You need to optimize for groups of users who share the same interest. That’s why the editorial question shouldn’t be “which articles generated Discover traffic?” but “which entities and topics consistently connect our journalism to an audience?”
This becomes most important at the interest-matching stage. Discover cannot associate content with an interest it cannot identify. If a user’s profile shows a strong interest in a person, team, company or television series, your article can only be matched to that interest if the headline and first paragraph clearly identify those entities. “TikToker lands a major opportunity” is generic. “Charli D’Amelio signs a deal with Netflix” identifies a specific entity. The user’s interest graph is useful to you only if your content’s entity graph is equally clear and structured.
Two consequences follow:
- Discover depends on authority built in Search. Our data shows that a domain Google trusts on a particular topic in Search is more easily associated with that topic in Discover as well. This is also why publishers with a strong Search foundation tend to maintain their Discover exposure better during updates. In our experience, publishers with a traffic distribution of 98% Search and 2% Discover generally have two to three times more growth potential on Discover, simply because they have never adapted their content to this surface.
- The feedback loop is where personalization becomes particularly strict. Card dismissals, short sessions and “Not interested” taps modify the user model against your domain, one reader at a time. Clickbait can work today and cost you dearly tomorrow.
We cannot see how much each of these sources weighs. But you don’t need to know the individual weights to act on the underlying principle: identify the entity clearly, deliver on the promise within the first screen and place the most important entities in the OG title (meaning the value of the page’s
og:titletag, i.e. the Open Graph title, Ed.). In our tests, this is the field Discover uses first.
For a small publisher that lacks the bargaining power of a major media group, blocking AI crawlers may seem like the only way to protect its content. But could that make the publisher less visible in AI-generated answers? How should it evaluate this trade-off?
Smaller publishers should keep two decisions separate that continue to be confused: blocking AI training and blocking AI distribution. They use different controls and involve very different costs.
Blocking Google-Extended in the robots.txt file limits the ability to use your content for Gemini training and grounding. Google says this does not affect your presence in Search results, rankings, AI Overviews, Top Stories or Discover. My recommendation for most publishers, large or small, is to block Google-Extended unless they have a commercial agreement with Google or a strategic reason to provide content to Gemini. You lose nothing in terms of visibility.
Blocking distribution is a different matter. Opting out of generative AI features through Search Console removes your links from AI Overviews, AI Mode and generative Discover features. For a small publisher, this is not a form of protection but of invisibility, because larger competitors will gladly occupy the citation spaces you leave open. The same logic applies to blocking crawlers from ChatGPT, Perplexity and other services: you protect your textual content, but you disappear from the answers. And to be clear: blocking these third-party crawlers has no effect on AI Overviews, which use Googlebot.
I understand why publishers object. The economics are uncomfortable: you pay to produce original reporting and AI systems use it to answer users without returning an equivalent volume of traffic. But being invisible does not protect a business. It simply makes the loss less visible. I mapped all these controls, explaining what each one actually blocks and what it costs, in our guide Top Stories in AI Overviews and the impact of blocking Google’s AI, because much of the confusion comes from treating them as a single switch.
The evaluation should therefore start with three questions:
- How much of my current visibility and traffic comes from these features? You need to measure this before making a decision and do so at multiple levels: understand whether AI platforms associate your brand with your topics, whether they cite you and whether that citation sends anyone to your site. I outlined this model in the guide AI visibility tracking for news publishers: the 3 layers that matter. For most small publishers, traffic generated by citations in AI Overviews is still limited today, but exposure through generative Discover features may not be. The level where the long-term value for the brand lies, however, is the association between the brand and its topics.
- Is my content really the kind of content AI systems can replace? If you publish standardized summaries, blocking AI changes nothing because AI can synthesize the same information using ten other sites. If you publish original local reporting, AI systems need you, and the citation is the moment when your brand becomes visible to a reader who otherwise might never have found you.
- Can I protect what matters without blocking everything? Using data-nosnippet on proprietary paragraphs, putting a paywall around deeper analysis and keeping distribution open for the summary layer is a more defensible position than a total block.
Small publishers also need to be honest about the reality ahead, because it will become harder, not easier. Search clicks are declining while impressions are increasing, AI Mode uses 13 to 39 sources per answer instead of the traditional ten links, and the consolidation analysts had been discussing is now visible in the acquisition market as well. Digiday reported last December that media M&A deals are being put on hold because of AI Overviews, that many publishers have lost a third or more of their traffic since the launch of AI summaries, and that owners uncertain about the future of their businesses are being advised to sell immediately at lower valuations. Independent sites such as The Planet D have already shut down, while Stereogum lost 70% of its advertising revenue in a year.
Blocking a crawler changes none of this. What makes the difference is scale and independence: producing more content and more original content than the newsroom produces today; being willing to join a larger organization or collaborate with one when remaining independent means downsizing; and building revenue streams that do not depend on search at all, through newsletters, events, subscriptions, licensing and direct relationships with the audience.
The goals I outlined in our Publisher Survival Playbook are a starting point: Google at 50% or less of total traffic, a direct relationship with most regular readers and at least 40% of revenue coming from sources other than advertising. A small publisher that reaches these numbers can afford to decide whether to block crawlers based on its principles. One that doesn’t reach them cannot afford to make that decision.
Small publishers don’t win by hiding. They win by becoming the source that has something others don’t and making sure AI systems know their name when that information is used.
If a publisher can choose not to appear in Google’s AI features but risks losing the visibility of its Top Stories inside AI Overviews as well, what choice should it make? Is it better to protect content from AI or leverage that visibility to continue reaching users?
Take the visibility, with two exceptions.
The reason is what the data shows about how these positions work. In the results we monitor, Top Stories inside an AI Overview is not a second position. Google displays the carousel either inside the AIO or as a standalone module beneath it, not both. When the carousel is integrated into the AIO, the headline, image, brand and link appear high on the page, inside the dominant search feature. When you opt out, our interpretation of Google’s documentation is that you lose that embedded position. Google has not, however, documented whether a standalone carousel reappears for the same query or whether the embedded carousel simply continues to be shown with other publishers. Being eligible for standalone Top Stories does not guarantee that Google will actually display one.
A publisher that opts out is therefore trading a known, highly visible position for an undocumented alternative. For almost everyone, that is a bad trade.
The two exceptions:
- a publisher that has a licensing agreement or legal strategy in which exclusion is a negotiating tool rather than the end goal. That is a business decision, and the loss of visibility is the price of the negotiation.
- a publisher whose exposure is concentrated in categories where embedded Top Stories in AIOs appear very rarely. In our data, health and science are close to zero. In these cases, the opt-out carries a lower Top Stories cost, although Discover exposure is still affected.
For everyone else, the path should be: measure how frequently your Top Stories placements are embedded in AIOs, broken down by section; measure how much exposure you receive from generative Discover features; if you still want to test the opt-out, use the inheritance between primary and secondary properties available in Search Console to apply it to a single section rather than the entire domain. Then compare the results.
A controlled test on a single section is worth more than an opinion about the entire site.
In your article, you distinguish between Consensus and Information Gain. If content still needs to cover the information Google and users expect to find, how can it add enough Information Gain to become genuinely different and more useful than the results already present in the SERP?
Consensus is the admission ticket. Information Gain is what gets you the ranking and the citation. Most publishers stop at the admission ticket.
A guide to Paris needs to talk about the Eiffel Tower, otherwise it feels incomplete. That’s consensus and it is necessary to build trust. But if a reader opens the first result, isn’t satisfied, opens your article as the second result and finds the same facts presented in a different order, Google has learned that your page adds nothing. That’s the loss, and AI makes it even more obvious because it compresses ten pieces of content based on the same information into a single paragraph and cites two of them.
Let’s stay with the Paris guide. Every result on the first page says the same things: visit the Eiffel Tower, book tickets online, the view is better at sunset, avoid the queue by going early. That’s consensus and you still need to provide this information. Information Gain is what comes next.
It’s the paragraph written by a journalist who spent a week at the entrances and explains which entrance had the shortest queue at 9 a.m., 1 p.m. and 7 p.m. It’s a table showing how much the same ticket actually cost, purchased on the same day through the official website, two resellers and a hotel concierge. It’s a statement from a guide who has worked at the Eiffel Tower for twenty years explaining which floor tourists regret skipping. It’s the local perspective: the bench on the other side of the Seine where Parisians go to watch the Tower light up, including the name of the street. It’s the data showing what people who read your previous Paris guide then asked in the comments, together with the answer to that question.
None of this requires a month of work. But these are elements the other nine articles don’t have, and they represent the part an AI summary cannot generate independently. That’s therefore the part it needs to cite.
The practical test I propose to publishers is simple: what is in this article that a reader could not already find in the first three ranking results? If the answer is “nothing”, then it is standardized content.
Information Gain does not require a six-month investigation. Small elements of novelty are enough when everyone else has a novelty level of zero:
- a quote from a named expert whom nobody else has interviewed;
- original data: a number that you own and others don’t, from your archive, audience or market;
- a perspective the other ten articles have ignored, often the local or industry-specific consequence of a national story;
- a question that the SERP reveals and nobody answers. People Also Ask is a map of information gaps;
- an update the others have not yet incorporated. Freshness on the same URL becomes Information Gain when the situation has changed.
The workflow matters as much as the intent. Before publishing, you need to compare the draft with what is ranking, not with what the author remembers about the topic. That’s exactly what we developed with NewzDash Article Optimizer: it analyzes the competitors ranking best for the article’s primary query and identifies missing entities, angles and reader questions, allowing the editor to see the gap before publication instead of discovering it three days later in the traffic report.
Cover the baseline information to build trust. Then add what only you can add. Information Gain is not saying the same thing using different words. That’s rewriting. It means adding something to the information ecosystem that wasn’t there before. And AI has made this even more valuable, not less, because producing an average article has become extremely easy.
What is the most common mistake publishers and SEO professionals make today when trying to create content with genuinely distinctive value?
Confusing “more” with “different”.
The typical response to the idea that “our content needs to stand out” is to produce a longer article: more sections, more subheads, more FAQs, more words. That’s comprehensiveness, and comprehensiveness is exactly what an AI system can produce on its own. For years, SEO taught us to read the top ten results, cover everything they cover and publish something bigger. The result is the average of the SERP, and today a model can generate that same SERP average in a few seconds. A 2,500-word deep dive that repeats what five other articles say is not distinctive. It’s consensus, developed at greater length.
The question I would ask every editor is this: if this article disappeared from the internet tomorrow, what information would disappear with it? If the answer is “none”, there is no Information Gain, however long the piece may be.
The second version of the same mistake is pretending to be different. You add an “expert opinion” without saying who the expert is, write “studies show” without citing a study, or insert a strong opinion that nobody in the newsroom actually holds. Readers notice, and increasingly systems notice too, because behind that statement there is no identifiable entity, source or numerical data.
The solution requires discipline, not a particular format. Before assigning an article, you need to establish what its distinctive element is. If the team cannot define it in one sentence (“we have the county data”, “we spoke to the surgeon”, “we have followed this company for six years and know what has changed”), then the piece is a rewrite and should be treated as such: fast, short, published to capture the query and not expected to generate traffic for three days.
Publishers that work this way effectively operate at two speeds. Speed allows them to win the first two or three hours on Top Stories and Live. Depth delivers results over the following three days through organic search, Discover and AI citations. The mistake is trying to assign both jobs to a single article.
Looking at the next few years, what signals do you think will allow Google to distinguish between content that is simply comprehensive and content that provides genuine Information Gain?
Google already has most of these signals. What changes is the weight they will carry when AI can generate complete and exhaustive content for free.
These are the elements that, in my view, will carry more weight:
- Entities and original claims. Google’s work on information gain, which dates back to its own patents, concerns measuring what a document adds compared with documents the user has already seen. As Google’s Knowledge Graph and the models behind AI Overviews become better at extracting information and claims, a page that introduces a new fact, a new named source or a new numerical data point will be clearly different from a page that repeats information already known. This is a signal about individual claims, not word count.
- Attribution to the original source. Who published a piece of information first, who was cited by others, who was linked to as the story spread. In news there is a real timeline: publisher A reports a fact at 10:02, publisher B picks it up at 10:20, a hundred sites have summarized it by noon and the web has reached consensus. But the information originated at 10:02, and Google and AI systems have every incentive to become better at identifying that moment. Original journalism leaves a trail of citations. Rewrites don’t.
- Author and brand entity strength on a specific topic. I’m not talking about a byline alone, but an entity: a recognizable person whom Google can connect to years of articles and reporting within a particular category. Our News Topic Authority data shows that authors who dominate a niche belong to a small, consistent group. Models that use sources to construct AI answers will rely on the same principle.
- Post-click satisfaction. Discover already has explicit signals in this area: follows, hidden content, “not interested” clicks, along with dismissals and very short sessions, clearly contribute to determining what the user sees next. Complete and exhaustive content that nobody finishes tells Google something. The same is true of a shorter article that users read to the end and share.
- Second-click behavior. This is a hypothesis, but I expect it could become an important signal: if a user opens an AI Overview, then opens one of the results and stops searching, that result probably added something. If they instead open three results and still return to run another search, none of those pieces of content truly satisfied their need.
- Multimodal originality. Original images, original video and original data visualizations that do not exist elsewhere on the web. Discover already places great importance on image quality, and AI systems can identify a reused wire-service photo as easily as they can recognize republished text.
I don’t expect length, keyword density or the number of subheads to carry significant weight, because these are exactly the elements a model can produce at virtually zero cost. For twenty years, we optimized content so machines could understand it. Now machines can understand mediocre content perfectly well. The advantage will come from giving them something genuinely worth understanding.
For publishers, the direction is clear. The next two years will reward newsrooms capable of proving they were the original source: the citation, the number, the document, the photograph. This is primarily a journalism problem before it is an SEO problem, and the SEO team’s job will be to make sure Google and AI systems can see and recognize that evidence.
Search, Discover and AI are no longer separate disciplines. They are three layers of distribution for the same journalism. That’s the theme we’re discussing with NewzDash publishers, and it will also be the theme of the conversation with a community from more than 50 countries at NESS in October 2026.
Beyond Google, publishers are now being cited, or ignored, by ChatGPT, Perplexity, Claude, Gemini and AI Mode. How should a publisher measure its visibility within these AI platforms, and can it actually influence it?
Start by accepting that keyword-based SEO tools and habits cannot simply be transferred into this new context. There is no first-party search volume for a prompt. Two people asking the same question to the same AI assistant on the same day can receive different answers with different sources. And most queries that matter to a news publisher are not prompts such as “best laptop 2026”, which remain stable for months. They are questions about a person, product, company or event that may exist for a week.
Measurement therefore needs to be built differently. I think about it in three levels, which represent the model behind the AI monitoring system we developed at GDdash:
- Association. When someone asks an AI platform about the topics on which you want to build authority, does your brand appear? Not just as a cited source, but as an entity that the model associates with that topic. A national broadcaster that never appears when someone asks Gemini about an election has a bigger problem than a missed citation.
- Citation. When the platform provides sources, are you among them? How often, in what position and against which competitors? This is the closest thing to a ranking and is the level where most tools stop.
- Traffic. Does the citation send anyone to the site? For almost every publisher we work with, referrals from ChatGPT and Perplexity are still negligible compared with Search and Discover. That will change, but a citation without a click is brand exposure and should be considered and reported as such, not as a traffic outcome.
Methodology matters as much as the measurement levels. You cannot monitor 50 prompts and call it a complete system. GDdash’s AI Visibility Tracking tools build prompt sets from four sources: the queries in the publisher’s Search Console, because long-tail queries phrased as questions are the ones users tend to paste into an assistant; the entities the publisher talks about most often on Discover; AI-generated prompts based on variants around the brand, competitors and generic searches for each topic; and prompts manually entered by the newsroom because they are considered relevant. We then run them repeatedly across different platforms, looking at trends rather than individual observations, because a single run tells you almost nothing.
Can you influence the result?
Yes, and in a less mysterious way than the industry suggests. Platforms can only cite what they can retrieve and understand, and their retrieval systems rely heavily on what search engines already rank, as well as content that is clearly structured, clearly attributed to a source and quick to retrieve.
Source selection is not identical to Google’s, but there is significant overlap. A publisher that earns a Top Stories position with an entity-rich headline and a fact placed within the first 100 words is giving these systems exactly the kind of material they can cite. A publisher that provides a distinctive fact, an identifiable expert or a number nobody else has reported will receive more citations, because a model cannot synthesize something only you have reported. There is no new AI-specific ranking factor to chase. There is authority, clarity and Information Gain, measured on a new surface.
If you had to advise a publisher today, what would you tell them to focus on over the next twelve months between Search, Google Discover and AI search?
Four things, in this order, and the order matters. Most teams are approaching them in the opposite direction.
First, protect Search, because it remains the baseline. Google Search and Discover send news publishers far more traffic than all AI platforms combined, by a wide margin, and that will still be true twelve months from now.
The fundamentals that allow you to win Top Stories during the first hours of a story — indexing within minutes, entity-driven headlines, a fact on the first screen, structured pages and fast templates — are the same fundamentals that support citations in AI Overviews. Teams losing ground today are, in most cases, not losing to AI. They are losing to a competitor who publishes faster on the same story. News SEO operates in minutes, and a daily position report tells you what you have already lost.
Second, take Discover to its full potential. In our analysis of more than 400 publishers, the share of traffic coming from Google Discover rose from 37% to 67.5% in two years, while Web Search fell from 51% to 27%. Publishers with a traffic distribution of 98% Search and 2% Discover generally have two to three times more room for growth.
The solution isn’t to produce more content. It is to specifically package 5% to 10% of daily output for Discover: strong entities in the OG title, images at least 1,200 pixels wide, a clear “why now” and a category-based editorial strategy built around who is winning that category, not simply what worked the previous week.
And the result needs to be measured at category level, because the February 2026 Discover update and the entertainment category reset at the end of last year demonstrated that a domain can do everything correctly and still move with its category.
Third, start measuring AI visibility immediately, even before it becomes a traffic line item. Not because ChatGPT will replace Google next year, but because the publisher with twelve months of data on association and citation trends when traffic arrives will know what to do, while the one starting from zero will be guessing. The cost of measurement is limited. The cost of being late is building a strategy around a competitor’s screenshot.
Fourth, and this is what also determines the other three: invest in journalism rather than rewrites. Every trend I’ve described in this interview points in the same direction. AI Overviews compress consensus-based content. Discover penalizes clickbait through its feedback system. AI platforms cite the source that reported the fact first. A newsroom producing ten rewrites a day is optimizing for a world that is disappearing. A newsroom producing two quick rewrites and three articles containing a citation, number or document that nobody else has is optimizing for the world that is coming.
What I would not put among the priorities is the next acronym. Search, Discover and AI are three layers of distribution for the same journalism. The winners over the next twelve months will be publishers that treat them that way, measure what is actually happening at each level and adapt faster than the competing newsroom.
What really matters in AI Search (no tricks)
So, what did I tell you? This really was a packed interview, full of incredibly useful insights.
If you work in publishing or digital in general, there is one concept you need to remember: Search, Discover and AI are not three separate battles. They are three layers of distribution for the same content, and if you treat them as separate silos, you are already losing ground.
My takeaway from listening to John Shehata is this: the most common mistake today is chasing the latest AI trick while ignoring the fundamentals. A clear headline with the right entities, a real fact in the opening lines, a piece of data that nobody else has.
This works across Google, Discover and AI citations because the machine is always looking for the same thing: something genuinely worth picking up and citing. The real competitive advantage in 2026 is called Information Gain, not recycled and rehashed content. Keep that firmly in mind!
A huge thank you to John for the clarity with which he shared data, methods and figures that normally remain confined to discussions among specialists and industry insiders.
And are you ready to hear from our next guest?
I’ll see you next week, right here on SEO Confidential, with another leading name in international SEO, GEO and AI Search.
#avantitutta
