What can AI mention tracking actually measure?
These tools measure the direction your brand is moving and how you sit against competitors. That distinction matters more here than it ever did in search, because a keyword ranking sits still between checks while an AI answer redraws itself every time someone asks. You came from a world where position 4 meant position 4 until Google changed it. This is not that world.
The gap is already visible in how the market behaves. Only 14% of marketers track AI citations while 43% call AI search optimization a core 2026 strategy, according to Digital Applied's share-of-voice framework. The work has run ahead of the measurement.
So the honest read is this. If you buy one of these tools expecting the clean, auditable number your rank tracker gave you, you'll either distrust the whole category the first time the figure jumps, or worse, you'll defend a number to leadership that can't survive a re-run. Buy it for the trend line and the competitive gap. That's what holds up.
Why do AI answers change every time you ask?
The same prompt returns different brands on repeat runs because these models generate text by sampling from a probability distribution rather than reading from a fixed table. Temperature and top-p introduce variation, as does the model's own token-by-token guessing, so a single answer reflects only one draw from the distribution. Even the machine's floating-point math shifts slightly between servers.
Researchers at Stanford and UC Berkeley put numbers on how unstable this gets. In one longitudinal study, some models changed their final recommendation on an identical, repeated prompt up to 40% of the time, as reported in the Journal of General Internal Medicine. Same input, different answer, four times out of ten.
Here's what that means for a report with one number in it. If a mention count can flip 40% on a re-run, then any figure you pulled from a single pass is one coin toss dressed up as a measurement. More passes are the fix, which is a point I'll come back to when we get to reporting.
How much does personalization skew results?
A tracking tool almost never sees what your actual customer sees, because AI answers bend to account history, location, and whether the person is logged in. The tool runs in a clean, neutral environment. Your buyer runs inside months of their own search behavior.
Google spells this out for its own systems. Personalized recommendations can reorder results based on a signed-in user's Search history, and location from past sessions refines what surfaces. The same logic drives the AI answer layer.
So when a tool reports you're absent from a category answer, read it as absent in a vacuum. A logged-in customer in your service area, with your brand in their history, will see you fine. Neutral-environment data provides a floor for evaluating the lived experience.
How do model updates reset your data?
Vendors ship model updates that move outputs overnight, which breaks your historical trend without any warning label. A sudden drop in visibility reflects a model change underneath you while your content retains its ground.
The Stanford and Berkeley team documented how sharp these swings can be. GPT-4's accuracy on one classification task fell from 84% to 51.1% between the March and June 2023 versions, per DeepLearning.AI's summary of the study. Nobody told users the behavior had shifted.
The practical lesson is to annotate your timeline the day a major model version lands. Because if you don't mark it, you'll spend a Monday explaining a cliff in the chart that has nothing to do with the work your team did. A discontinuity in the data is not always a discontinuity in your performance.