A session in partnership with Page 2 Podcast
Some SEO fixes feel too obvious to test.
Fix broken breadcrumb schema. Rewrite a basic title tag. Add an internal link. Put the obvious keyword in the obvious place. These are the kinds of recommendations that have filled technical audits and SEO tickets for as long as I have been in the industry.
The problem is that obvious is not the same as proven.
That was the theme of this conversation with Jon Clark and Joe DeVita. We talked about SEO experimentation, the origins of SearchPilot, what makes a good test, why enterprise teams need evidence rather than instinct, and how LLM traffic changes the measurement problem without removing the need for Google.=
The awkward truth is that best practices can lose traffic.
I have seen breadcrumb schema fixes go negative. I have seen title tag changes swing wildly in both directions. I have seen perfectly sensible ideas turn into inconclusive tests because the change was too small, too constrained, or blocked by an internal team before it touched the part of the page that mattered.
That does not mean experience is useless. It means experience should produce hypotheses, not certainty.
˙✧˖ AI-written summary
Below is an AI-assisted summary of the webinar conversation. This is not a word-for-word transcript but is included to help you find the key parts of the conversation.
Why best practice is not proof
The podcast opened with a simple challenge: what if SEO best practices were costing teams traffic?
That sounds dramatic, but it captures a real problem in enterprise SEO.
A lot of SEO work is shipped because it sounds obviously correct. A technical audit finds missing markup. A title tag looks under-optimised. A product page lacks a piece of internal linking. A template could include more keywords. The recommendation feels sensible, so it gets added to the roadmap.
The missing step is evidence.
Will's point was that SEO is full of high-confidence recommendations that behave differently once they are tested. Some win. Some lose. Some do nothing. Some depend heavily on the site, template, vertical, competitors, SERP layout, user behaviour, and timing.
That is why SearchPilot exists.
The goal is not to prove that SEO expertise is useless. The goal is to turn expertise into testable hypotheses. A strong SEO team should still have instincts, pattern recognition, and experience. But those instincts should feed an experimentation programme, not replace one.
SearchPilot has written about this distinction in its guide to what SEO split testing is, which explains why changing a subset of pages and comparing them with a control group gives teams a cleaner read than before-and-after reporting.
How SearchPilot came out of Distilled
A large part of the conversation covered the business story behind SearchPilot.
Before SearchPilot, Will spent 15 years building Distilled into a global search agency. At its peak, Distilled had around 60 employees, offices in London, Seattle and New York, the SearchLove conference series, DistilledU, a strong content engine, and hundreds of clients.
But that period also came with strain.
Will described the earlier 2010s as a time when the business hit its first real growth wall. After years of top-line growth, Distilled had its first year where revenue fell compared with the previous year. For an entrepreneur who had internalised growth as a sign that everything was working, that was difficult.
Out of that period came a renewed focus on R&D.
The early version of SearchPilot did not begin as a standalone software company. It began as software inside an agency: a way to make Distilled better, more differentiated, and more effective for clients.
The team was looking for ways software could improve SEO consulting. Tom Anthony, now CTO of SearchPilot, was part of that early R&D team. What eventually became SearchPilot started life as DistilledODN, a software-enabled agency capability.
The clean break came later.
In 2019, Distilled began conversations with Brainlabs about an acquisition. During that process, it became clear that the agency and the software business should separate. Brainlabs wanted the profitable agency. Will, Duncan, and the team working on the software were excited by the product opportunity.
So the software was spun out.
Then 2020 happened.
The new company started life as an independent software business just before the pandemic, which was not exactly the calm reset anyone had in mind. But that spin-out forced a sharper question: what does SearchPilot do better than anyone else?
The answer became controlled SEO experimentation for large websites.
Why focus became the real business lesson
Jon and Joe also asked about the shift from running an agency to running a software company.
Will's answer was that focus became much more important.
Distilled had done many things at once: consulting, conferences, training, publishing, R&D, international expansion, and software. Some of that was exciting. Some of it was chaotic. Will described himself as someone who is naturally attracted to shiny new ideas, which made focus something he had to learn deliberately.
SearchPilot is different.
It is built around one core capability: helping enterprise teams run SEO tests at scale.
That focus affects product, sales, marketing, customer success, and even which customers SearchPilot serves best. The company has narrowed around large ecommerce, retail, marketplace, and travel sites where there are enough pages, enough traffic, and enough commercial value to make controlled testing worthwhile.
That is an important business lesson, but it is also an SEO lesson.
Focus matters in a testing programme too. Teams do not need endless ideas. They need the right ideas, prioritised well, built cleanly, tested with enough sensitivity, and interpreted in a way the business can use.
How SEO A/B testing actually works
A useful part of the conversation covered the difference between SEO A/B testing and traditional CRO testing.
In a CRO test, users are usually split into buckets. One user sees the control experience. Another sees the variant experience. That works because the goal is to measure how users behave.
SEO testing has a different problem: the search engine crawler is part of the system.
A team cannot simply bucket Googlebot into one experience and users into another without breaking the test or creating cloaking problems. If Googlebot only sees one version, then the search engine is not really being split.
SearchPilot solves this by splitting pages, not users.
A group of pages is selected as the control group. A comparable group is selected as the variant group. The variant pages get the change. The control pages do not. The pages are balanced statistically so that historical traffic patterns are as similar as possible between the two groups.
The change is made server-side and is visible to everyone who visits the variant pages: users, Googlebot, and LLM-related crawlers.
That matters.
There is no cloaking. There is no separate crawler-only version. The page experience is consistent for every visitor to that page.
This page-level approach is central to SearchPilot's explanation of how SEO A/B testing works, and it is why large ecommerce and travel sites are such a strong fit for the platform.
What makes a site testable
The podcast also covered the practical requirements for running SEO tests.
The key ingredients are traffic and pages.
Will mentioned a rough rule of thumb: around 30,000 organic sessions per month, or about 1,000 per day, distributed across a useful number of pages. That does not mean every site needs to hit that exact number, but it gives a sense of the scale required.
Teams also need enough pages to create control and variant groups.
Two pages are not enough. Dozens can sometimes work. Hundreds or thousands are better. This is why large ecommerce and travel sites are a strong fit: they often have templates with many pages receiving meaningful organic traffic.
Conversions are important for the business, but they are not always the best statistical success metric for the test itself. Revenue and conversions can be noisy because they are affected by discounts, competitor pricing, promotions, seasonality, macro conditions, and many other factors.
Organic traffic is often a cleaner primary metric for detecting the effect of a page change. The business can then translate that expected traffic impact into revenue using its own models.
SearchPilot's page on what it does explains this wider platform approach: server-side SEO tests across large groups of pages, with results that can be tied back to commercial impact.
What makes a strong SEO hypothesis
One of the most important parts of the conversation was the difference between a test idea and a hypothesis.
A test idea might be:
"Add FAQ content to these pages."
A hypothesis explains why that change should affect search performance:
"Adding concise FAQ content to category pages will help them rank for additional long-tail queries and improve relevance for existing queries, increasing organic sessions to the tested page group."
That difference matters because weak hypotheses lead to weak tests.
Some tests are not bad because the idea is wrong. They are bad because the expected mechanism is unclear, or because the change is too small to produce a measurable result. Will mentioned the possibility that AI could help SearchPilot build a "this test is not going to do anything" detector: not a tool that predicts winners and losers, but one that identifies underpowered ideas that are unlikely to move the needle.
He gave alt attributes as an example.
There are many good reasons to add alt attributes to images: accessibility, usability, compliance, and sometimes image search. But Will said SearchPilot has not seen evidence that changing alt attributes moves standard organic SEO traffic in either direction. That does not make alt attributes unimportant. It means they may be a weak candidate for a traffic-focused SEO A/B test.
For teams trying to improve this part of their process, SearchPilot's guide on how to write a strong SEO hypothesis is a natural next read.
Why test backlogs grow faster than teams can run them
A common concern from new customers is whether they will run out of test ideas.
SearchPilot's experience is the opposite.
Large SEO teams usually have more ideas than they can test. Those ideas come from in-house SEOs, agencies, product teams, content teams, merchandising teams, technical audits, competitor analysis, internal politics, and old recommendations waiting for engineering resource.
The challenge is not idea generation.
The challenge is prioritisation.
Will said SearchPilot often helps customers judge which tests are most likely to move the needle and which are easiest to build. The goal is to find the overlap: changes that are both commercially meaningful and practical to implement.
The larger blocker is often organisational.
Some teams cannot test above-the-fold on PDPs because the product team owns that space. Some cannot change templates because engineering resource is scarce. Some cannot test certain content modules because brand or merchandising owns them. Some get stuck in sign-off loops.
When test velocity is low or win rate is disappointing, the issue is not always the quality of the SEO ideas. It may be that the team is only allowed to test low-impact parts of the page.
That is why enterprise SEO experimentation is as much an operating model as a tool.
Why title tags are still dangerous
Title tags came up as one of the clearest examples of why "simple" SEO changes can be powerful and risky.
Will said title tag tests are among the most likely to produce very large positive results and among the most likely to produce very large negative results.
That is because title tags affect two things at once.
They can influence rankings.
They can also influence click-through rate.
Even though Google rewrites some titles, it does not rewrite all of them. The title a team writes still often contributes to the search result snippet, which means it affects how compelling the result is to searchers.
Will reflected on old SEO recommendations where teams would do good keyword research and then recommend adding those keywords into title tags. Looking back, he sees that as a much higher-variance recommendation than many SEOs treated it at the time.
The paid search comparison is helpful.
Paid search analysts have long known that one advert can dramatically outperform another because the copy is better, not because it contains a different number of keywords. Organic snippets work the same way. The wording matters.
Will mentioned a title tag test where the first version was catastrophically bad, with a decline of more than 20% in organic traffic. But the underlying idea was still useful. The team iterated and found a version that produced a large positive result.
That is the lesson.
The idea was not "title tags work" or "title tags do not work."
The lesson was that implementation matters.
What the breadcrumb schema test taught us
Jon raised one of SearchPilot's more memorable examples: a breadcrumb schema fix that reduced traffic.
This is exactly the kind of thing SEOs usually ship without much hesitation. If breadcrumb markup is broken, fix it. That sounds like technical SEO housekeeping.
But the test produced a negative result.
The lesson is not that breadcrumb schema is bad. The lesson is that search result changes can affect user behaviour in ways that are hard to predict.
Structured data can change the way a result appears. A technically cleaner search result may not be a more clickable search result. In some cases, the old appearance may have been more compelling, or the rich result may have changed the user's expectation in a way that reduced clicks.
SearchPilot has written about this in its breadcrumb test case study, SEO test: How important is breadcrumb markup for SEO?. It is also discussed in SearchPilot's post on DIY SEO testing pitfalls, which explains why before-and-after comparisons can mislead teams when they are not using a proper control group.
This is the uncomfortable part of SEO testing.
The thing that looks right in an audit can still lose traffic.
What carries across sites, and what does not
Jon asked whether SearchPilot's large body of test data produces transferable rules.
Will's answer was nuanced.
Some learning does transfer. SearchPilot can learn which kinds of changes are likely to move the needle. Title tags, for example, are high-variance and powerful. Tiny changes to things that have never moved the needle are less promising.
But the direction of impact is much less transferable.
A title tag change may be a big winner on one site and a big loser on another. A breadcrumb change may help one template and hurt another. A content block may improve one category and do nothing elsewhere.
That is because every site is different.
The competitive set is different. The SERP layout is different. The user expectation is different. The brand strength is different. The template is different. The internal linking is different. The algorithm has changed. Competitors may copy the advantage six months later.
This is why SearchPilot often retests similar ideas over time.
The goal is not to build a secret master list of SEO answers. The goal is to build a system for finding the right answer in a specific context.
Why LLM traffic changes the measurement problem
The conversation then turned to AI search and LLM traffic.
Will's view was that SearchPilot's focus on ecommerce and travel makes the LLM measurement problem more tractable than it is in some other verticals.
In ecommerce, there is still an action at the end of the journey.
ChatGPT cannot secretly ship a pair of trainers. A retailer still fulfils the order. In most cases today, the customer still reaches the retailer's website at some point, even if the early research happened inside an AI interface.
That gives ecommerce teams something to measure.
SearchPilot is now measuring traffic from sources like ChatGPT, Perplexity, and other LLM-powered platforms. In many cases, that traffic is bundled into an LLM or AI referral segment so teams can see whether a change moved that audience.
But Will's preferred view is still net impact.
Google traffic is still much larger for most customers. A change that dramatically improves LLM referrals but slightly hurts Google could be bad for the business if the Google loss is larger in absolute terms.
This is the same argument SearchPilot has made in the Omio GEO A/B testing story: one Omio test increased LLM traffic by +18%, while another showed that a GEO-positive change would likely have hurt Google organic sessions by -6.5%. Without measuring both, the team could have made the wrong rollout decision.
Why net impact matters more than channel wins
The phrase "net impact" is one of the most important ideas in the conversation.
It is tempting to celebrate a win in a new channel because the channel is exciting. AI search is new. ChatGPT traffic is growing. Leadership is paying attention. A dashboard showing LLM referrals rising feels like progress.
But ecommerce teams cannot ignore the existing business.
For many enterprise sites, Google organic traffic is still 10 or 100 times larger than LLM referral traffic. That means even a small Google decline can outweigh a large percentage gain in LLM referrals.
The right question is not:
"Did LLM traffic go up?"
The better question is:
"What happened to total qualified traffic and business performance?"
This is why SearchPilot's GEO A/B Testing measures AI and Google performance together. SEO and GEO can overlap, but they are not always identical. A test needs to show the trade-off.
That is especially important as teams start testing AI-friendly content, product summaries, structured key takeaways, review summaries, and richer product detail. Many of those changes could plausibly help AI retrieval. Some may also help SEO. Some may not.
The only safe answer is to measure.
How LLMs retrieve fresh information
Will explained that LLM testing mainly targets the retrieval side of AI search, not the model's training data.
A model's training data is fixed until the next training run. Optimising for what GPT-6 might learn in the future is not a very practical marketing strategy. It is slow, uncertain, and almost impossible to attribute.
Retrieval is different.
When someone asks a search-like question in ChatGPT or another AI system, the system often grounds its answer in live search results or fresh web data. It does this because the model needs information that changes: prices, stock status, offers, reviews, product availability, competitor dynamics, and new content.
That is why LLM traffic can be tested.
If a page change affects how a site is retrieved, interpreted, or cited in those fresh-answer workflows, the impact can appear in LLM referral traffic.
This connects back to classic SEO. Search engines have spent decades learning that freshness matters. Query deserves freshness, faster crawling, and live indexing were all attempts to solve the same problem: the web changes, and the answer needs to reflect that.
AI has not removed that problem. It has made it more visible.
SearchPilot's article LLMs do not rank anything. So what are you optimizing for? goes deeper on this point: LLMs are not ranking pages in the old sense, but their answers often depend on retrieval systems that are still connected to search visibility.
Why crawl data helps explain test timing
Jon asked about log files and whether SearchPilot uses them to validate or invalidate tests.
Will explained that SearchPilot does not technically work with raw server log files in the traditional sense. Log files are generated as a side effect of visitors and crawlers hitting a website, and at enterprise scale they can be huge, fragmented, and difficult to access.
But SearchPilot does look at crawl data.
Historically, crawl data has not been used as the final success metric for tests. A test is not considered a win simply because Googlebot crawled more pages.
But crawl data is very useful for timing and interpretation.
If a test is currently neutral, one important question is whether Google has actually seen the change yet. If only 10% of the test section has been recrawled, the lack of movement is not very meaningful. The page change cannot affect search performance until the search engine has discovered it.
So crawl data can help teams decide whether to wait.
For example, if a title tag test has not moved yet, but most pages have not been recrawled, the team should probably not rush to interpret the result. Once recrawl reaches a higher level, the test becomes more informative.
SearchPilot's article on using Googlebot tracking data in SEO split test analysis expands on this idea: crawl data can help explain what is influencing a test result and whether the test has had enough time to be meaningful.
What UGC and reviews mean for AI search
The conversation also covered user-generated content and reviews.
LLMs cannot physically try a product. They cannot wear the trainers, use the umbrella, sit on the sofa, or take a suitcase on holiday. They rely on what humans and trusted sources have said.
That makes reviews, product experience, expert commentary, and first-hand information more important.
SearchPilot is not trying to test every off-site influence, such as social media, PR, or influencer content. The platform is focused on what customers can change on their own websites.
But on-site reviews and UGC are part of that.
Teams can test whether reviews are present in the HTML rather than hidden behind JavaScript. They can test prominence. They can test summaries. They can test whether additional product detail, manufacturer information, or review-derived content changes performance.
Will said some of SearchPilot's more successful LLM-focused tests have involved giving models more information. The source of that information could be manufacturer data, product specifications, reviews, or other product details.
The key point is that an LLM does not get bored reading detail. A human may not want an endless wall of specifications, but a model can use that detail to answer specific questions.
That creates a design challenge.
The page still needs to work for humans. But the underlying information also needs to be accessible and useful to machines.
Why AI makes experimentation more important
One of the most useful parts of the conversation was Will's answer to the question of whether AI changes SEO testing.
In one sense, the answer is yes.
SearchPilot is measuring LLM referral traffic. Teams are testing changes that might improve visibility in ChatGPT, Perplexity, and similar systems. Fan-out queries, retrieval, freshness, and LLM interpretation all create new hypotheses.
In another sense, the answer is that AI reinforces the original reason SearchPilot exists.
Search was already opaque. Google was already a black box. Machine learning already shaped search results long before AI Overviews and AI Mode. The only thing that has changed is the level of uncertainty and executive attention.
A checklist mindset was already weak.
AI makes it weaker.
This is why the same discipline matters: form a hypothesis, make a controlled change, measure the result, and decide whether to roll it out.
SearchPilot's AI content testing roadmap is a good example of this approach. The question is not whether AI-generated content is good or bad in the abstract. The question is whether a specific AI-assisted change helps a specific site section without creating a downside elsewhere.
The business lesson: do the typing
The podcast ended with a more personal discussion about focus, metrics, and leadership.
Will described himself as more of an optimiser than a target-setter. Some people are naturally motivated by clear numerical targets, personal bests, and scoreboards. Will's answer was more values-based: he looks for whether he feels good about the work, whether it was valuable, and whether it aligned with what he is trying to achieve.
One phrase stood out: "I've gotta do some typing."
For knowledge work, thinking is not enough. Planning is not enough. Ideas are not enough. At some point, the work has to move from the head to the page, the product, the test, the customer, or the team.
That applies to SEO experimentation too.
It is easy to debate best practices. It is easy to collect ideas. It is easy to read case studies and argue about what should happen.
The work is running the test.
What teams should take away
The main takeaway from this conversation is that SEO teams need to move from "this should help" to "we tested this."
That does not mean testing every tiny change. It does not mean ignoring expertise. It does not mean reducing SEO to statistics.
It means recognising that enterprise SEO is full of uncertainty.
A best practice can lose.
A title tag can be a huge win or a huge loss.
A schema fix can change the search result in a way users do not prefer.
A test idea can be too weak to matter.
A GEO change can help LLM referrals and hurt Google traffic.
A page may not move because Google has not recrawled it yet.
A strong SEO programme needs a way to separate opinion from evidence.
That requires:
- clear hypotheses
- enough traffic and pages
- balanced control and variant groups
- server-side implementation
- crawl visibility
- business interpretation
- cross-functional permission
- enough test velocity to learn faster
The awkward truth is that "best practice" is often just a hypothesis that has not been tested on your site yet.
Put Search in Control Mode with SearchPilot
This conversation with Jon and Joe came back to one of SearchPilot's core beliefs: search is too important to run on instinct alone.
For enterprise ecommerce, retail, marketplace, and travel sites, organic search is often one of the biggest channels in the business. It is also one of the hardest to attribute cleanly, one of the most exposed to platform changes, and now one of the channels most affected by AI search.
SearchPilot helps enterprise teams make SEO and GEO testable. It runs controlled experiments across product pages, category pages, templates, navigation, internal linking, content, structured data, Merchant Center surfaces, and AI-influenced journeys, then helps teams understand what moved and what the result means.
The goal is not to replace SEO judgement.
The goal is to give good SEO teams evidence.
That is how teams stop relying on before-and-after comparisons, stop shipping best practices on faith, and stop treating AI search as a separate panic channel.
Search is your biggest channel and least understood. Take it out of react mode. Put it in control mode.