Buying journeys are better modelled as conversations, not prompts, and randomness at every step means that prompt tracking misses a lot of the picture.
It might sound obvious, but you can't model buying journeys as single prompts to an AI agent. Just like in the old-fashioned search days, real buying journeys involve multiple steps of discovery, interest, and research. Unlike in that era, when each of those steps might have resulted in a click to your website, big chunks of this buying journey now happen within walled gardens where we can't see them.
This fact doesn't make these discovery platforms less valuable. If anything, it adds to their power, but it does make it difficult for marketers to join the dots between the data and real business impact.
I'm seeing too many AI visibility reports that are essentially meaningless percentages, disconnected from business impact and untethered from causality. There's no way of telling whether the numbers have moved because of things you did, things the platforms did, the growth of consumers' use of AI, things your competitors did, or seasonality. At the same time, it's impossible to tell how much it's worth to your business to move that visibility number from 43% to 45%.
But that's not the only problem I have with prompt visibility tracking. Each of these could be an article in its own right:
- I am sceptical of the value of building KPIs out of synthetic prompt data which are guesses about what people ask
- I am sceptical about the provenance of volume data, whether panellists really understand what it's going to be used for, and how representative it therefore is of the real market
- I am sceptical of the value of prompts that are dropped into a fresh context without history, memory, or tools and connectors, because it fundamentally fails to model how real users are using AI
As I find myself saying a lot these days, "Everything is a unique search now."
But it's worse than that. It's not only that every paragraph-length prompt I type is most likely unique in the history of prompts. It's also only one in a sequence of prompts that I might do on a buying journey, and it will be shaped by the history and context that the AI holds about me.
Does this matter, though? I wanted to get a sense of how much variation is likely to be introduced into the recommendations that an AI might make as conversations branch off in different directions, so I set up an experiment.
Experiment design: modelling a conversation
My goal was not necessarily to simulate real-world conversations, but rather to build a toy model to see the sensitivity to initial conditions and the kind of divergence and probabilistic behaviour that we might expect even under quite controlled conditions without searcher history and account context.
I set the experiment up as follows for each kind of shopping goal:
- I described the goal in a high-level single sentence. For example, “helping a middle-aged man pick running shoes for his first half marathon”
- Claude then fleshed this out into a detailed persona with a made-up backstory, attributes, and demographic information, held back as context for generating the conversation, but not supplied verbatim to the AI model in the research process
- A small local LLM with no other tools or internet access was then given the context of the persona and access to a single public frontier model to do its research: in this case, Gemini via the API
- The small local LLM went back and forth with Gemini, starting with a simple query and building up, depending on the results it received, to ask further questions or clarify its requirements in more detail until it was happy that it had reached a good recommendation
- Claude oversaw the whole thing and analysed the back-and-forth conversation, including the final recommendation received by the small local model.

I ran steps 3, 4, and 5 twelve times each for two different shopping journeys.
What I learned
The biggest takeaway is the extreme lack of stability in the final recommendations of what to buy, even though the task and the persona were the same across each independent run.
In particular, the final product chosen was most often NOT the top recommendation on the initial prompt.
I saw a lot of drift between the recommendations in the first response and those at the end of the conversation. Different runs ended up recommending a lot of different products: four distinct final brands across 12 [running shoe] runs, and nine across 12 [garden furniture] runs. Once you add conversational dynamics (path dependence, plus the probabilistic nature of each LLM response), it's clear that we're dealing with a much more variable environment than we can model by tracking a stable set of prompts over time. Maybe that load of academic words is easier to visualise with a diagram:

Check out visualisations of the outcomes here.
Example 1: running shoes
Brooks was the first recommendation in half the runs, so a prompt tracker would have Brooks comfortably winning this query. But only one of those six conversations ended with Brooks. Saucony was recommended first in four runs and was chosen in none of them, yet it was the most common final choice overall (five runs), mostly from conversations that had opened with Brooks. Only two runs out of 12 ended where they started (the coloured ribbons).

Example 2: garden furniture
This is a much more fragmented category, and it shows: seven different brands or retailers were recommended first, and nine different ones were chosen at the end. Wayfair was the most common opening recommendation (three runs) and won none of them. Beliani, the most common final choice (three runs), didn't appear in a single opening answer. Again, only two of the 12 runs finished where they started. Nine of the 12 conversations also hit the turn limit and had to be forced to a decision. A bigger, more considered purchase with more constraints (patio size, seating, assembly, delivery timing) takes longer to resolve, and that gives the path even more room to wander.

So what should marketers do?
None of this means AI assistants don't matter for discovery. Quite the opposite: in both experiments, the assistant was deciding which brands made the shortlist and clearly influenced which one won. What it does mean is that a visibility score built from the first answer to a fixed prompt is measuring the wrong thing. In the running shoes example, it would have told you Brooks was winning, when Saucony was the brand that actually got chosen most often. And that was with a shopper that had no memory, no history, and no other context. Real people's conversations are likely to diverge more than this, not less.
First, stop treating prompt visibility percentages as KPIs. At best, they're a rough signal of which brands are in the conversation, not of whether you're winning it.
Second, pay attention to the whole journey, not just the opening answer. This is why we are still paying a lot of attention to visits and purchases when measuring AI discovery performance for retail.
Measuring the real business impact of website changes is exactly what we built SearchPilot to do. If you'd like to talk about what that looks like for your site, get in touch.