I like making bets about the future. Not because I think I am especially good at predicting it, but because writing down exactly what you think will happen forces you to be more precise.
Kevin Indig and I made one of these bets a little over a year ago in June 2025. His side is that by 31 December 2030, ChatGPT will have more visits than Google, measured by Similarweb over a rolling 30 day period. I took the other side.
Kevin admitted in this conversation that his side is not looking especially good right now. Google moved much faster on AI than he expected. But the interesting part of the bet was never really who gets bragging rights in 2030.
The interesting question is what has to happen to user behaviour, Google, ChatGPT, ecommerce and the rest of the web for either of us to win.
Useful links from the conversation
- AI Halftime Report: H1 2026 - Kevin's wider research on AI search, measurement, consumer behaviour, trust, agentic workflows, and attribution.
- SEO and GEO are not the same: Omio's GEO A/B tests - a SearchPilot customer story showing why an AI visibility win can still hide a Google organic downside.
- Users behave differently in AI Overviews - Kevin's research into how people interact with AI-powered search experiences.
- The consensus gap - research into how ChatGPT, Gemini, and Perplexity differ in the sources they cite.
- When is a click not a click? - Will's take on why fewer clicks can still mean more valuable visits as AI does more of the research before the click.
- The alpha is not LLM monitoring - Kevin's argument for treating AI visibility tracking more like directional research than traditional rank tracking.
˙✧˖ AI-written summary
Below is an AI-assisted summary of the webinar conversation. This is not a word-for-word transcript but is included to help you find the key parts of the conversation.
The Google vs ChatGPT bet
The bet started with a chart Kevin shared while ChatGPT usage was growing very quickly. He projected the traffic curves for Google and ChatGPT forward and argued that they could cross before the end of 2030.
Kevin's original logic followed classic disruption theory. A large incumbent has an established product and business model. A challenger arrives with a different technology and a rapidly growing audience. The incumbent struggles to adapt because the new product threatens the thing that made it successful.
Will took the opposite view. His bet was partly about status quo bias: five years feels like a very long time in AI, but changing the behaviour of billions of people can take much longer. More importantly, he thought generative AI might strengthen Google's product rather than force Google to abandon its existing business.
That distinction now looks important. Kevin said during the session that he would probably make the same bet again with the information he had at the time, but he also acknowledged that his side is currently behind.
Google moved faster than expected
Kevin said there were two developments he had underestimated. The first was how quickly Google would respond. AI Overviews, AI Mode and Gemini let Google put much of the new AI experience directly inside products people already use.
The second was OpenAI's execution. Kevin felt OpenAI had enormous momentum but spread its attention across several initiatives while Google caught up quickly on the core discovery experience.
Will's explanation came back to Clayton Christensen's distinction between disruptive and sustaining innovation. If generative AI improves Google's existing search business, Google's incentives point in the same direction as the technology. In that case, the incumbent can move surprisingly quickly.
Kevin compared the situation with Facebook copying Snapchat Stories. Facebook did not need every Snapchat user to switch. It could make the new format available to its existing audience and reduce the need for new users to go somewhere else. Google's distribution gives it a similar advantage.
Search is becoming about behaviour as well as intent
Kevin thinks one of the more interesting shifts is that AI lets search move beyond the immediate query.
Classic Google Search has an unusually strong intent signal: the user explicitly types what they want. Social feeds work differently. They infer interests from behaviour, such as what someone watches, clicks, likes or ignores.
AI can combine the two. A system can read a detailed prompt while also knowing more about the person asking it, their previous conversations, their past purchases or the wider context around the request. Kevin expects that combination of intent and user behaviour to become more important.
Will agreed that personalisation will grow, although he pointed out that a detailed conversational prompt can make intent even stronger. If someone has spent 30 seconds explaining exactly what they need, trying to push them towards something unrelated may still be difficult.
AI traffic can be smaller and more valuable
This matters because the meaning of a click is changing.
Will said SearchPilot is already seeing higher conversion rates from AI referral traffic across large retail customers. Kevin said he has seen similar patterns, including higher average order values. The likely explanation is that the AI has already done part of the research and comparison before the visitor reaches the retailer.
That means fewer clicks do not automatically mean less commercial value. A user who has compared products, refined their requirements and arrived ready to buy is a very different visitor from someone who clicked the first blue link at the start of their research.
This is why raw referral traffic is an incomplete way to judge AI discovery. SearchPilot has written separately about how AI traffic shows up in analytics, but the more important question is what those visitors do once they arrive.
Ecommerce and media have different problems
Will drew an important distinction between businesses that sell information and businesses that sell something a customer ultimately has to buy.
For media, publishing and affiliate businesses, AI can absorb much more of the product itself. If the user's goal is to get information, a sufficiently good summary may remove the need to click through. That makes dependence on referral traffic increasingly risky.
Retail and travel are different. ChatGPT can research the shoes, compare them and recommend them, but somebody still has to sell and ship the shoes. The customer still needs a merchant, brand, retailer or marketplace somewhere in the transaction.
The number of clicks may fall because the model effectively opens and reads the research tabs on the shopper's behalf. But the remaining click near the point of purchase can become extremely valuable. That is one reason SearchPilot is increasingly interested in experiments that measure AI referral traffic alongside Google rather than treating AI discovery as a separate visibility exercise.
Agentic commerce may look more like Apple Pay
Kevin and Will were both sceptical of the most extreme version of agentic shopping: someone vaguely says they need new shoes, then a bot independently chooses a pair, buys them and has them delivered.
Will's nearer-term version looks more like "fancy Apple Pay". The person decides what they want, then the agent handles the tedious part. It knows the preferred delivery option, which retailer has the loyalty account, which payment method to use, and perhaps which merchant is most reliable.
Kevin's view is somewhat more bullish on native AI checkout. His user behaviour research suggests people already place considerable trust in AI-generated shopping shortlists. He thinks that once platforms can qualify merchants, display trust signals and complete checkout inside the experience, many users may have little reason to leave.
Will thinks trust will slow that change. People still want detailed product information, high resolution images, returns policies, price confidence and reassurance about the merchant. The disagreement is mainly about timing, not direction.
Users behave differently inside AI search
Kevin has run multiple studies where normal US adults complete tasks inside AI Overviews, AI Mode, ChatGPT and other AI experiences while researchers record what they do.
The results suggest that AI interfaces create behaviour that looks quite different from classic search. In one shopping study, users rarely challenged or fact-checked AI Mode's shortlist. Kevin said around 75% selected the first result on average, showing a strong preference for the top recommendation.
Trust was one of the few things powerful enough to override that position. When users recognised and trusted a brand further down the shortlist, they frequently chose it instead. That makes brand familiarity and trust more important than a simple "rank number one" model might suggest.
Kevin also found that the outbound clicks in that study happened when users were ready to transact. That is another reason ecommerce teams need to understand the full journey rather than measuring only how often a link appears.
The consensus gap
Kevin's research on the consensus gap challenges the idea that "AI visibility" is one thing.
He looked at the sources cited by ChatGPT, Gemini and Perplexity and found very little overlap. Only around 2% of URLs in his study were cited across all three engines. Platform differences were larger than differences between individual models on the same platform.
That means a brand can perform well in one AI environment and poorly in another. The recommendation itself can also change because the systems are grounding their answers in different sources.
Will connected this with SearchPilot's Omio GEO A/B testing story, where a single website change produced different effects in ChatGPT and Google. It is increasingly difficult to talk about "GEO" as though it were one homogeneous channel.
Why blended averages hide the real story
A lot of AI visibility dashboards collapse these differences into one score.
Kevin thinks that is too blunt. An aggregate number may look healthy while hiding weak performance on the engine that matters most to a particular audience. The opposite can also happen: one weak platform can drag down a blended score even when the brand is performing well where its customers actually search.
His recommendation is to start gathering evidence about which engines customers use. That could include self-reported attribution during signup or checkout, referral data, customer research and other first-party signals.
The goal is to move from "What is our AI visibility score?" to more useful questions: How visible are we in ChatGPT? In Gemini? In Perplexity? Which of those platforms matters to our customers? What sources are driving the answers there?
What LLM monitoring is actually useful for
Kevin's article The alpha is not LLM monitoring argues that the current prompt tracking model needs to evolve.
Most tools take a sample of prompts, run them repeatedly and aggregate the answers into visibility or citation metrics. Kevin thinks teams should separate engines and models more carefully, and he questions whether daily tracking adds much value when the answers themselves are highly variable.
His broader framing is that prompt tracking should look more like polling than rank tracking. Instead of pretending a synthetic prompt is a deterministic keyword position, teams can treat a set of model responses as a sample that tells them something about how the system currently understands the brand.
Will described essentially the same idea as market research. A focus group is artificial too. Nobody mistakes ten participants in a room for the entire market. But it can still reveal patterns, language and problems worth investigating.
Prompt tracking has a leadership use case
Will added one caution for SEOs who are understandably sceptical about AI visibility dashboards: large companies are buying them very quickly.
He joked that many prompt tracking platforms seem to have stronger product-market fit with leadership reporting than with proving business impact. But that leadership demand is still a real signal. Boards and executives want to know what is happening with AI.
Kevin agreed that this is often why CMOs care. The board is asking the question, so the organisation needs some way to answer it. The problem comes when the visibility dashboard becomes the end of the analysis rather than the start.
The useful next step is application. What should the team change because of what it has learned? Can that change be measured? Does it improve performance in the engines customers actually use? Does it help AI discovery without hurting Google? SearchPilot's GEO A/B Testing is designed around those questions.
What search teams should do next
The conversation did not end with a prediction that every company should move budget from Google to ChatGPT, or the reverse.
The practical work is more specific. Understand how customers actually use different AI environments. Separate performance by platform instead of relying on one blended visibility number. Treat prompt monitoring as directional research. Measure what happens after the referral. For ecommerce, pay particular attention to trust and purchase behaviour.
Teams should also be cautious about treating today's interface as permanent. Search is becoming more personalised. AI platforms disagree with one another. Checkout may move closer to the AI experience. Google is absorbing AI into products it already owns.
That uncertainty is the argument for testing rather than waiting for someone to publish the definitive AI search playbook. SearchPilot's work with Omio has already shown that a change can help LLM traffic while hurting Google organic performance. The net impact matters more than winning one dashboard.
Put GEO in Control Mode with SearchPilot
Kevin and Will may still disagree about who wins their 2030 bet, but the practical implication for search teams is the same.
Trying to predict whether Google, ChatGPT, Gemini, Perplexity or another platform will dominate five years from now is interesting. Building your strategy around that prediction is much riskier.
SearchPilot helps enterprise teams make SEO and GEO testable. Teams can run controlled experiments across category pages, product pages, content, internal linking, structured data and other high-impact site sections, then measure Google and AI referral performance together.
Stop trying to predict the future. Experiment to discover it.