Return to Articles 12 mins read

Is GEO Working? How to Get Beyond Prompt Tracking

Posted August 14, 2026 by Will Critchlow

I wanted to do something a little different with this webinar.

Usually, these sessions are conversations. This one is me laying out where we have got to with GEO testing, what we have learned from real experiments, and why I think ecommerce teams need to move past guesswork quickly.

The reason is simple: there is a lot of noise right now.

Every day, someone on LinkedIn shares a new checklist. And they often contradict yesterday’s. Use structured data. Structured data does nothing. Chunk your content. Chunking is useless. Write summaries. Do not write summaries. Track prompts. Prompt tracking is broken.

Leadership is asking questions faster than the industry is producing reliable answers. "Are we ready?", "Are we doing the right things?", "How do we go faster?". These are good questions, but they need better answers than another dashboard of AI visibility scores.

That is what this session is about: how to stop guessing and start testing.

 

˙✧˖ AI-written summary

Below is an AI-assisted summary of the webinar conversation. This is not a word-for-word transcript but is included to help you find the key parts of the conversation.

Why prompt tracking is not enough

Will opened the session by naming the problem many teams are feeling: AI search and discovery has created a new measurement market but hasn’t connected the data to causality or business impact.

There are many tools that track brand mentions, citations, AI visibility, and share of voice via synthetic prompts across ChatGPT, Perplexity, Gemini, AI Overviews, AI Mode, and other surfaces. Some of that data is useful. But it does not answer the bigger question leadership is asking:

Is this working?

Prompt tracking can tell a team whether it appeared in a sample of AI answers for a sample of prompts. It can help with debugging. It can show early warning signs. It can help teams understand how a brand is described in certain contexts.

But it cannot prove that a website change made the business better.

Prompts are personal. They are often long, specific, and shaped by previous conversations. The same user might ask a follow-up question that changes the whole context. Different users may get different answers. The universe of possible prompts is effectively infinite.

That is why Will described this as a "search volume one" world.

The problem is not that prompt tracking is useless. The problem is that prompt tracking is being asked to do too much. It is not the same as measuring impact. It is not the same as proving that a change should be rolled out across a large ecommerce site.

For teams trying to understand how AI visits appear after the click, SearchPilot's guide to how AI traffic shows up in analytics is a useful companion piece.

Why ecommerce is such a good testing ground

The session focused heavily on ecommerce because ecommerce has a particular advantage in the AI discovery era: people still need the manufacturer and retailer to make and ship products.

ChatGPT is not going to ship the shoes or manufacture the product. It may help people research, compare, and decide, but the retailer, marketplace, travel site, or brand still has a role to play.

That makes ecommerce different from some media and publishing models, where the content itself may be more directly substituted by an AI summary.

For retail and transactional websites, the risk is real, but so is the opportunity. Will's view is that more people will be buying more things online in the coming years, and many of those buying journeys will be shaped by some form of organic discovery. The interface may look like a mix of search engine and chatbot. The discovery process may be more conversational. But the commercial question is familiar:

Will the customer buy from you?

This is why product and category pages are so important. Large ecommerce sites usually have scalable templates: PDPs, PLPs, category pages, internal search pages, faceted pages, buying guides, and related content blocks. Those templates can be tested.

That makes ecommerce a practical testing ground for GEO. Teams can change something across a controlled set of pages, compare performance against a control group, and measure what actually happened.

SearchPilot's GEO A/B Testing is built around that same principle: test changes on real page groups, measure AI and Google performance together, and avoid relying on opinion.

How LLMs find fresh product information

Will explained that AI discovery has two broad information sources.

The first is training data. This is the information the model absorbed during training. It shapes the model's understanding of language, entities, relationships, associations, and brand context. But it is largely fixed until the next training run.

That is a slow timescale. A team cannot go to leadership with a strategy that amounts to "wait for the next model and hope it likes us more."

The second source is retrieval.

When a user asks a question, the model may bring in fresh information during the interaction. This is often described as retrieval-augmented generation, or RAG. Will described it as the model opening fifteen or fifty tabs in the background, reading across the web, and bringing the user a synthesised answer.

For ecommerce, this is essential.

A model cannot rely only on old training data to answer questions about:

  • current stock levels
  • today's price
  • active discounts
  • latest reviews
  • delivery options
  • product availability
  • new launches
  • updated product details
  • local availability

Product recommendations need freshness.

This is where traditional search and AI discovery reconnect. AI systems need up-to-date information, and that often means retrieving pages, feeds, or search results from the live web.

That is also why SearchPilot's work around Merchant Center Testing matters for ecommerce teams. Product feeds, structured data, PDPs, pricing, availability, and product attributes can all become part of how machines understand and recommend products.

What fan-out queries change

The next important concept was fan-out queries.

A user may type one long prompt into an AI system, but the model may break that task into many background searches. It may search for product comparisons, reviews, pricing, availability, best options for a use case, brand reputation, delivery details, and other supporting information.

The user sees one answer.

Behind that answer, there may have been many searches.

That changes how teams should think about optimisation. The old model was often keyword-first: what keyword are we targeting, where do we rank, and what does the search result look like?

In AI discovery, the hidden fan-out queries may be where the real retrieval happens.

Teams usually cannot see all of those fan-out queries. They may get clues but they cannot treat the process as a clean list of keywords.

That is one reason testing becomes more important.

A team can make a change, then measure whether that change improved LLM referrals, Google organic traffic, or the net business outcome. It does not need perfect visibility into every hidden query to measure the effect.

This connects closely to SearchPilot's article LLMs do not rank anything. So what are you optimizing for?, which explains why old ranking-factor thinking can become misleading in AI search.

How to write a GEO hypothesis

Will compared SEO hypotheses with GEO hypotheses.

In traditional SEO, a successful test usually works through one of three mechanisms:

  • targeting new keywords
  • improving rankings for existing keywords
  • changing the search result appearance so more searchers click through

GEO has analogues, but the language changes.

A GEO hypothesis might aim to:

  • target new fan-out queries
  • improve visibility for existing fan-out queries
  • influence the summary returned by an LLM
  • make a page, product, or brand easier for the model to recommend

The fourth mechanism is especially interesting because it feels a little like conversion rate optimisation, but the "converter" is partly the machine. The question becomes: has the page provided the information the model needs to confidently recommend the product?

That could involve product detail, comparison language, reviews, freshness, structured data, key features, FAQs, delivery information, stock information, or buying guidance.

A weak GEO hypothesis says:

"This might help AI visibility."

A stronger GEO hypothesis says:

"Adding clearer product suitability information to PDPs may help models retrieve and recommend these products for more specific fan-out queries, while also improving confidence in the AI-generated summary."

That gives the team something testable.

SearchPilot's AI content testing roadmap is relevant here because many AI-era content changes need to be framed as hypotheses, not treated as obvious best practice.

Where GEO testing happens

For large ecommerce sites, GEO testing happens on the same kinds of scalable surfaces that SEO testing already uses.

That includes:

  • product detail pages
  • product listing pages
  • category templates
  • buying guide modules
  • comparison content
  • FAQs
  • review summaries
  • key feature summaries
  • internal linking modules
  • structured data
  • freshness indicators
  • product feed-aligned content
  • availability and delivery information

The mechanics look similar to SEO A/B testing. A team makes a change to a variant group of pages, compares performance with a control group, and measures the result.

What changes is the journey being measured.

In traditional search, a user might open several tabs, compare sources, read reviews, check products, and then come to the site. Much of that research was visible across a set of searches and visits.

In AI discovery, more of that research may happen inside the conversation. The model reads, compares, summarises, and narrows options before the user arrives. The site may only see the final click.

That makes the click more valuable in some cases, but harder to interpret.

SearchPilot has written separately about this in When is a click not a click?, because a lower number of clicks can still represent more qualified visitors when more research happens before the visit.

Why GEO and SEO can disagree

One of the most important points in the session was that GEO and SEO can disagree.

Many GEO changes could plausibly help SEO too. More useful content, better structure, fresher product information, clearer summaries, stronger internal links, and better structured data can all have SEO hypotheses attached to them.

But that does not mean every GEO-positive change is SEO-positive.

Will explained that most practical GEO work today still reaches AI systems through search-related retrieval. That creates overlap with SEO. But overlap is not the same as sameness.

A change can help an LLM understand and summarise a page while hurting Google organic performance. A change can make a page richer for AI retrieval while making it bloated, duplicative, or less effective in traditional search.

This is where the danger of single-channel measurement becomes obvious.

A team could look only at LLM referrals, see a positive result, and roll the change out. But if Google organic traffic falls by more in absolute terms, the business loses.

That is the bigger risk with guessing what works in GEO: a visible win in one channel can hide a larger loss elsewhere.

SearchPilot's GEO A/B Testing is designed to measure these effects together, so teams can see the net impact rather than celebrating one metric in isolation.

What the Omio test taught us

The clearest example in the session came from SearchPilot's work with Omio.

Will explained that Omio partnered with SearchPilot to share early GEO testing results with the wider industry. The key lesson was not only that GEO can be tested. It was that GEO and SEO do not always move together.

SearchPilot has now published the Omio GEO A/B testing story. In one test, adding brand USPs increased LLM traffic by +18%. In another, adding structured key takeaways performed positively for LLM-driven traffic, but would likely have hurt Google organic sessions by -6.5%, so Omio chose not to roll it out and developed follow-up iterations instead.

That is the practical value of testing.

The AI result looked positive on its own. The business result was not.

Without measuring Google organic performance at the same time, the team could have rolled out a net-negative change.

This is the strongest argument against treating GEO as a checklist. A tactic can be directionally plausible and still wrong for a specific site, page type, or business goal.

What prompt tracking can actually tell you

Prompt tracking came up again later in the session as one of the most common tools teams are adopting.

Will said that almost everyone he speaks to is using one of the large prompt tracking or AI visibility tools. He also hears almost everyone say some version of: "We do not quite know whether we trust the data, or what to do with it."

That does not make the tools useless.

The better comparison is rank tracking in traditional SEO.

Rank tracking is useful for debugging. It can help diagnose problems. It can show movement. It can surface early warning signs. It can help teams understand where visibility may be changing.

But rank tracking is not the same as business impact.

Prompt tracking has the same issue, with extra complications.

Teams do not know the full set of prompts users are typing. Many prompts are unique. Answers are personalised. The model may use memory, previous conversations, location, and other context. A brand may appear for a prompt in one run and not another. A dashboard can only sample a small portion of reality.

It can help teams form hypotheses. It can help teams notice issues. It can help teams explain some AI visibility patterns. But it should not be the main evidence that a GEO programme is working.

How to answer leadership

One of the strongest themes in the webinar was the sudden growth in executive attention.

Will shared a quote from an SEO leader at a public company: "I've had more questions from leadership in the last 6 months than in the preceding 10 years."

A mature response to leadership should explain:

  • what the team is testing
  • which templates or site sections are in scope
  • what hypotheses are being prioritizes
  • how AI and Google performance are measured together
  • what the team has learned so far
  • what has been ruled out
  • how test velocity will increase
  • what resources are needed to move faster

That kind of answer gives leadership something more useful than reassurance. It gives them evidence, cadence, and a path to learning.

SearchPilot's core SEO experimentation platform was built for this kind of operating model: controlled experiments, measurable impact, and a way to explain what changed to product, engineering, marketing, and finance.

What ecommerce teams should do now

The session ended with a clear message: teams need to build the capability to learn faster.

Nobody knows exactly what AI discovery will look like in one, two, or three years. The interfaces will change. The models will change. The role of links will change. AI referrals will change. User behaviour will change.

The answer is not to predict everything correctly.

The answer is to build a testing programme that can adapt.

This is the difference between reacting to AI search and building a system for it.

Put Search in Control Mode with SearchPilot

This webinar came back to the same idea that sits at the heart of SearchPilot's positioning: search is often your biggest channel and the least understood.

AI discovery makes that more true, not less.

There are more surfaces, more uncertainty, more leadership questions, and more tactics that sound convincing but may not work on a specific site. The way through is not another checklist. It is controlled testing.

SearchPilot helps enterprise teams make SEO and GEO testable. We run controlled experiments across product pages, category pages, templates, navigation, internal linking, content, structured data, Merchant Center surfaces, and AI-influenced journeys, then help teams understand what moved and what the result means.

The Omio story shows why this matters. A GEO-positive result can hide a Google organic downside. A test can stop a net-negative rollout before it harms performance. It can also show where AI discovery changes are genuinely worth scaling.

That is the answer leadership needs.

Not "we are tracking prompts."

Not "we followed the latest GEO checklist."

But: "We are testing what works on our site, measuring AI and Google together, and learning faster every month."

Search is your biggest channel and least understood. Take it out of react mode. Put it in control mode.

Sign up to receive the results of two of our most surprising SEO experiments every month