What can we still truly measure in the age of AI Overviews, AI Mode and LLMs? We talked about it with Semrush’s Owned Media lead
In this new episode of SEO Confidential we interviewed Nick Eubanks, one of the top figures at Semrush.
Today Nick leads the Owned Media division at Semrush, where he coordinates an international team focused on growing the company’s revenue and supporting a decisive strategic transformation: the shift from traditional SEO analysis tools to new solutions dedicated to visibility within the AI ecosystem.
In this conversation we tackled the most pressing issues of the moment, including:
how much can we trust the tools that measure visibility in AI answers and what limitations does AI Mode introduce?
We also discussed how generative models are changing the very meaning of “ranking”, especially now that transparency, traceability and reporting accuracy are becoming increasingly difficult to guarantee.
As you will see, we asked the questions many avoid, challenged a few certainties and pushed back on some of the industry’s most deeply rooted beliefs, trying to speak also on behalf of the readers who are most skeptical and doubtful about the usefulness of the new analysis and monitoring tools.
Nick didn’t dodge anything: he answered point by point, with great clarity and rigor.
But enough talk, enjoy the interview.

“The fact that models are opaque doesn’t mean they aren’t measurable”, said Semrush’s Nick Eubanks
Here we are, Nick. First of all I’d like to ask you: with the growth of AI Overviews, AI Mode and AI-powered answer engines, what will the future role of SEO tools be?
Will they remain essential for performance analysis and monitoring, or do they risk losing relevance in an ecosystem where traditional metrics are becoming less visible and measurable?
SEO tools will continue to matter, but their function will change: from analyzing rankings and traffic they will move to measuring visibility within AI systems, narrative impact and attribution across increasingly hybrid surfaces.
Traditional metrics such as rankings and organic traffic remain useful, even though they are no longer enough. Discovery increasingly happens through AI-driven interfaces, such as AI Overviews and conversational models, which filter content and influence choices without necessarily generating a click.
Tools will therefore need to introduce new capabilities: showing how content is presented in AI-generated answer panels, measuring brand presence in the narratives produced by the models, and connecting exposure on AI surfaces to subsequent behaviors, such as brand searches, direct visits and conversions. The goal is to attribute influence even when the interaction happens outside traditional paths.
If SEO tools keep focusing only on rankings, backlinks and classic reports, they risk losing relevance, because they miss the central dynamic: how AI selects content, which sources it cites and how that exposure changes brand perception.
In short: SEO platforms are not obsolete, they are transforming. They are becoming platforms for “Generative Engine Optimization (GEO)” just as much as for classic SEO.
Many tools promise to measure visibility in AI answers, but there are huge doubts about personalization, context and the technical limits of the models.
If these factors are not taken into account, what kind of reality do those numbers actually represent? What can you say to the skeptics who don’t believe in the usefulness of these systems?
When personalization and context are neglected, metrics become mere partial samples: they show what a profile similar to a certain user might see in a controlled situation, without capturing the full reality of the audience.
If variables such as location, history, device or model version are not considered, the data risks turning into noise. It offers an apparent visibility that doesn’t always match actual influence across all segments.
For those who raise doubts, it helps to make clear from the start that the goal is not absolute measurement but an indication of trends and relative share. The value comes from comparison: before and after a campaign, against competitors with the same set of prompts, or across different models.
The idea is close to TV audience ratings or survey panels: they don’t describe the entire population, yet they provide reliable indications. Likewise, AI visibility panels offer a plausible reference within an environment where complete data doesn’t exist.
Ignoring these measurements means giving up on understanding what happens in the “dark funnel” of AI-mediated discovery. Even if imperfect, those numbers remain useful for guiding analysis and decisions, as long as their limits are understood.
How can you prove the ROI of SEO to clients, especially when Google doesn’t make it available? Can tools like Semrush really fill this gap, or do they remain only an indirect indicator?
ROI can no longer be presented as a direct ratio between a single keyword and a precise economic result. On AI-mediated surfaces the reference becomes the overall effect, attribution modeling and the ability to understand how influence propagates.
A few tactics can make measurement more solid:
- Synthetic control groups: compare optimized pages or markets with similar non-optimized pages, observing over time the changes in organic visits, brand traffic and direct visits.
- Connecting visibility in AI answers to business results: when the brand appears in AI-generated answers or citations, brand searches and direct traffic often increase. The tools show the visibility changes, analytics show the final effects.
- Using platforms like Semrush to fill in the missing metrics: they provide data on share of voice, mentions, citations, topic difficulty and signals useful for estimating value. They are indirect indicators, but they represent the most reliable inputs in a context where the engines don’t provide complete information.
- Connecting visibility to conversion outcomes: greater exposure, across organic and AI surfaces, tends to correlate with higher conversion rates, lower acquisition costs and greater long-term value.
In short, tools like Semrush don’t offer a perfect ROI calculation, but they provide indispensable signals. Combined with modeling, analysis and reasoned assumptions, they make it possible to build credible assessments that clients can understand.
With the rise of personalization in AI Mode through the integration of user profiles and context, how do you plan to represent visibility and performance data realistically, considering that each user may see a different output?
It helps to recognize that every user can receive a different output, even though generative systems share a common “backbone” of retrieval and ranking. On this basis, reliable models can be built.
Visibility can be described on several levels:
- Global baseline: use a set of prompts with a neutral profile, with no history or specific traits, to identify the brand’s minimum presence.
- Segments based on different persona types: define representative profiles such as business buyer, consumer or healthcare professional; pair each prompt with the relevant context and check how visibility changes.
- Constant trend monitoring: keep prompts and test conditions unchanged, so you can observe variations over time. Even though real users see different results, it’s possible to measure percentage increases or drops in a controlled environment.
When presenting the data to stakeholders, it’s useful to clarify that it’s not a snapshot of the entire population, but a structured reading of the brand’s performance across various contexts and against competitors.
Segmentation also helps highlight relevant differences: if one profile shows the brand twice as often as another, that information has strategic value.
Full personalization means it won’t be possible to cover every individual experience, but it is still feasible to build a general-level map of trends that supports decisions and makes the behavior of generative systems more predictable.
AI Mode’s query fan-out fragments a single question into multiple subqueries and sources: are tools already able to map and measure the coverage of these subtopics to assess a site’s real authority?
In principle, and increasingly in practice, tools like Semrush are able to map the fan-out graph: seed query → subqueries → related entities/facets.
There are a few key points to analyze:
- For each initial intent, you need to identify the sub-questions that may be generated, using topic analysis, competitor research and AI prompts.
- You need to verify whether your content covers these subqueries and whether the site appears on AI surfaces through citations and mentions.
- It’s also worth checking visibility in traditional SERPs for the same topical nodes.
Measuring topical authority thus turns into a coverage assessment: in how many relevant subqueries the site is present, how often it is cited and how many elements of the topical map are actually covered.
This pushes tools beyond simple keyword counting or rankings, toward models based on intent coverage, visibility in AI answers and the assessment of the entity as an authoritative reference.
Beyond tracking inclusions, do you plan to integrate comparative metrics into your tools that clearly distinguish the performance of queries with AI Overviews, without AI Overviews, and with AI Overviews but without citations?
Yes, and this is a decisive differentiator. Segmentation should follow three distinct groups:
- Queries where no AI Overview is triggered.
- Queries where an AI Overview appears and the brand is cited.
- Queries where an AI Overview appears and the brand is not cited.
This breakdown makes it possible to analyze three scenarios:
- For queries without an AI Overview, you can observe performance in terms of traffic and rankings.
- For queries with an AI Overview where the brand is cited, you can assess the impact of that exposure on visibility and traffic.
- For queries with an active AI Overview but no citations, you identify the gap, and therefore an opportunity for action.
Comparing these segments helps identify strategic areas, such as high-volume queries with an active AI Overview from which the brand is absent. This makes it possible to set clear priorities for optimizing content and entities.
Tools should immediately highlight both the opportunities – high volume, AI Overview present and brand absent – and the risks, such as the case where the site ranks first in organic results but is excluded from the AI Overview.
So yes, integrating this kind of comparative segmentation is absolutely part of the roadmap for next-generation SEO tools.
Nick, isn’t it a bit paradoxical to want to “measure” a context, that of LLMs, which is by nature opaque, dynamic and constantly evolving? How much can brands trust this data?
Yes, it’s paradoxical, because the inner workings of generative models are proprietary and constantly evolving. But we already measure systems that are just as opaque (Google’s search algorithm, social algorithms, programmatic ad auctions).
Reliability comes from a few key elements:
- Consistent methodology: use stable prompt sets, defined personas and contexts, repeated over time.
- Reading trends instead of focusing on single measurements. Absolute values may vary, but the direction of change is a solid indicator.
- Cross-checking to look for correlations between AI visibility metrics and business results (brand search growth, increased direct traffic, conversions).
- Clarity about limits: communicate that the data represents sample-based estimates, not a complete picture of every possible user journey.
Brands can consider this data reliable if it is interpreted as a set of decision-making tools, not as absolute truths. Its value lies in answering questions such as:
- “Is our AI visibility growing more or less than our competitors’?”
- “Are we closing the gaps in our coverage of the topics where AI dominates?”
So the paradox is real, but that doesn’t mean you can’t act on it. You just have to do it with the right approach and the right discipline.
You have added Google AI Mode tracking to Semrush. Let me play devil’s advocate. How do you respond to those who say it’s hard to take seriously a tool that measures visibility in AI answers, given that no one can know how those models really work or how they choose their sources?
It’s true, we don’t have full access to the model’s decision function, but we can observe its outputs systematically.
We focus on the visible signals: which sources are cited, how often brands are mentioned, which prompts bring our content to the surface.
We run controlled prompt sets under fixed conditions (device, location, model version where possible) so we can compare like with like.
The fact that models are opaque doesn’t mean they aren’t measurable.
We measure outcomes, not the models’ inner workings, just as we have done for years with Google Search.
Measurement is about visibility and influence; it doesn’t guarantee inclusion. The tool helps you understand opportunities, gaps and trends.
So yes, it’s legitimate to question the mechanisms, but valid measurement is still absolutely possible and actionable if framed correctly.
You launched the first report for ChatGPT Shopping, but this feature is still in its infancy and not widely used by consumers. Are you offering brands a real advantage, or do you risk fueling hype that could fade quickly?
The realistic answer: it’s an early strategic advantage rather than a channel already capable of generating certain, predictable volumes.
There is evidence supporting this reading:
- OpenAI is clearly integrating commerce features into ChatGPT along with agentic commerce protocols, partnering with major retailers.
- Consumer behaviors already show the use of AI assistants for product recommendations and decision-making = a pre-purchase phase.
But there are also risks:
- Adoption is still embryonic; metrics/benchmarks aren’t as mature as those of classic search or e-commerce platforms.
- If brands invest heavily based on hype alone (ignoring the fundamentals), they may see limited returns.
So our goal is to develop the tools and expertise to be ready when this channel grows.
We’re not pretending this is your next 50% traffic channel today, but we are giving you visibility into a channel that could become important, and the knowledge gained will be applicable to all AI commerce surfaces.
So it’s not just hype, but an early stage. The advantage lies in early benchmarking, in learning cycles and in being ready when scale arrives.
LLMs tend to prefer brand websites over aggregators. Isn’t there a risk of redesigning a web that favors only big brands, erasing the role of the intermediaries that have always generated value?
The data suggests that generative systems show a preference for authoritative, entity-rich sources validated by experts and third parties. For example, the study “Generative Engine Optimization: How to Dominate AI Search” found that AI search favors qualified media outlets and authoritative third-party content over low-quality aggregated content.
This doesn’t automatically mean that only big brands will win, but it raises the bar for smaller publishers and intermediaries: they need to move from “volume aggregators” to “value creators”.
Intermediaries can stay competitive if they:
- offer original data, proprietary analysis, independent testing or expertise built around their own community;
- adopt a rigorous use of entities, structured data, authorship signals and references that help LLMs recognize their authority;
- own areas of vertical specialization where big brands lack depth or credibility.
Yes, there is a risk of market consolidation, as you say, but new possibilities can also emerge.
Generative systems reward differentiation, traceability of expertise and clarity.
The intermediaries that remain relevant will be those able to offer expert, well-defined contributions, not copies or rehashes of producers’ content.
The message for those working in publishing and content marketing is simple: you are not being excluded, you are entering a new phase. The environment rewards those who can become an authoritative intermediary, not those who merely produce volume.
Measurement in the AI era: metrics in crisis and new tools still to be deciphered
SEO is going through a phase shift that no professional can afford to ignore.
Generative models are redefining what it means to “be visible”, while analytics tools try to measure increasingly opaque surfaces, where transparency and causality are less and less guaranteed.
It’s new territory, full of enthusiasm and suspicion: some fear the data isn’t reliable, some see AI Mode as an earthquake for historical metrics, and some wonder whether LLMs are rewriting the web’s hierarchies in favor of big brands.
Throughout this interview we tried to represent your doubts and reservations as well, about the tools that measure something as seemingly elusive as AI visibility.
We have to admit it: Nick didn’t sidestep the questions and answered us frankly.
He acknowledged Semrush’s limits and potential, explained the opportunities and pointed out the direction with the clarity of someone living this transition from the inside.
It’s true, these tools aren’t perfect and don’t offer absolute precision, but they can provide estimates, projections and useful signals.
But in my opinion the real point is another: they do what they can despite the void left by Google.
Search Console should finally provide more accurate, more complete and less filtered data, instead of forcing us to work with partial and often ambiguous information.
Until that happens, these tools remain useful operational resources (despite their limits), because at least they try to fill a gap that shouldn’t exist.
A big thank you to Nick for his clarity and willingness to engage.
That’s all for today, see you next week with a new episode of SEO Confidential.
Our journey into the “backstage” of search continues!
