Stop Measuring AI Visibility Like a Vanity Metric
- Jonathan Bowman

- Aug 5
- 7 min read

There is a new report circulating in every marketing team right now, and it has a shiny name. AI visibility. The pitch goes like this: ChatGPT, Gemini, Perplexity, and Google's AI answers are becoming the front door to the internet, so you had better know how often they mention you, how they feel about you, and how much traffic they send. Fair enough. That is a real shift and I am not going to pretend it isn't.
But here is where it goes sideways. Most of the AI visibility dashboards I have seen are built to make you feel something, not to make you decide something. They count mentions. They score sentiment. They chart referral traffic from a handful of AI tools. And then everyone nods, because the line goes up and up feels like progress.
Let me ask the question nobody in the demo asks. Up compared to what? Toward what?
The metrics everyone is selling you
Walk into any pitch for an AI visibility tool and you will hear the same three headline numbers. Citations, meaning how often a model names or links to your brand. Sentiment, meaning whether the model describes you in warm words or cold ones. And LLM referral traffic, meaning the clicks that land on your site from an AI product.
None of these are useless. I want to be precise about that. They are early signals, and early signals have value when you are trying to understand a system nobody fully understands yet. The problem is not that these numbers exist. The problem is that people are treating them as the scoreboard when they are barely the weather report.
Counting how often a robot says your name is not a strategy. It is a horoscope with a chart attached.
Think about citations for a second. A model can name your brand in an answer that actively steers the buyer somewhere else. It can list you fourth out of five. It can mention you in a context that has nothing to do with what you sell. If I told you your company got named 400 times last month, your very next question should be: named how, to whom, and did any of those people ever become anything to us? If a tool cannot answer that, the 400 is decoration.
Why sentiment scores are mostly theater
Sentiment is the one that really gets me, because it sounds so sophisticated. We are measuring how the AI feels about your brand. Except the AI does not feel anything, and the score is a machine reading of machine output, which is a hall of mirrors dressed up as insight.
Here is the practical issue. Sentiment is volatile and it is prompt dependent. Ask the same model the same question two ways and you can flip the tone. Ask it on Tuesday and again after a model update and the whole baseline moves under your feet. You are trying to build a trend line on sand.
More importantly, positive sentiment and buying intent are not the same thing. I have watched brands with glowing, friendly AI descriptions generate nothing, and blunt, boring, purely factual mentions pull in real buyers. Warmth is not conversion. A model saying nice things about you is pleasant. It is not a lead.
That matters.
So when a dashboard shows me a sentiment gauge creeping from neutral to positive, my honest reaction is, okay, and? If I cannot draw a line from that needle to a single decision I would make differently, it is a mood ring. Nice to look at. I am not running a business on it.
The referral traffic trap
Now the one that feels the most legitimate. LLM referral traffic. Real humans, clicking real links, arriving on your real site from an AI tool. Surely that is the number that counts.
I understand the appeal. It looks like the old SEO game we already know how to play. But two things make it a trap if you lean on it too hard.
First, the volume is often tiny and the attribution is a mess. A lot of AI referrals do not show up cleanly in GA4. They get bucketed as direct, or they arrive stripped of the context that tells you why the person came. So you end up optimizing a number that is both small and half blind. That is a bad combination.
Second, and this is the bigger one, the entire point of these AI tools is to answer the question without the click. The model reads your content, synthesizes it, and hands the buyer a conclusion. The buyer never visits. If your only measure of AI visibility is the traffic that survives that process, you are measuring the leftovers. You are counting the people the machine failed to fully serve, and calling it success.
If you only measure the clicks that escape the AI, you are grading yourself on the answers the machine got wrong.
So no. Referral traffic is not the scoreboard either. It is a fragment of the story, and a shrinking one.
What I actually want to know
Let me flip this. Forget the tools for a minute and ask the question a founder actually cares about. When a real buyer, the kind who could spend money with us, goes to an AI tool and describes their problem, do we show up as a serious option? And when they dig deeper, does the model represent us accurately enough that we make their short list?
That is the whole game. Everything worth measuring hangs off it.
So the first thing I track is not brand mentions in general. It is presence in high intent, buying stage prompts. Not "what is content marketing" but "best content marketing agency for a B2B software company" and the fifty specific ways your actual buyers phrase their actual need. I want to know if we appear when someone is close to deciding, not whether we get name dropped in a definition.
Those are two completely different questions, and only one of them is connected to revenue.
Share of the answer, not share of voice
The metric I have come to trust most is something I think of as share of the answer. Not how often you are mentioned across the whole universe of prompts, but how often you appear, and how prominently, inside the specific set of questions your buyers actually ask when they are getting ready to spend.
Position inside the answer is part of this. Being the first, most confidently described option is worth a great deal more than being the afterthought at the bottom. We knew this in search. A model summarizing three vendors and describing you first, in specific and correct terms, is doing something close to a warm referral. Being listed last, vaguely, is closer to being ignored politely.
Accuracy belongs here too, and this is where sentiment should have been all along. I do not care whether the model is warm about us. I care whether it is right about us. Does it describe what we actually do, who we serve, and what makes us different, or is it repeating something two years stale and half wrong? A confident, accurate description in a buying moment is the asset. Fix accuracy and the warm words tend to follow anyway.
So the shortlist I actually pay attention to looks like this:
Presence in buying stage prompts, tracked against the specific questions your buyers ask, not generic category terms.
Position and prominence inside those answers, because first and specific beats last and vague every time.
Factual accuracy of how the model describes you, because a wrong answer in front of a ready buyer is worse than no answer.
Whether that visibility is actually feeding pipeline, which is the only number that ends the argument.
Connect it to pipeline or stop counting
That last point is the one I will die on. Every AI visibility metric has to be pressure tested against the same brutal question. Does this connect to a lead, to pipeline, to revenue? If it cannot, in some honest and traceable way, then it is a curiosity, not a KPI.
Here is how I try to build that bridge in the real world, because I know clean attribution from AI tools is hard right now. I do not rely on a single tracked link and pretend it tells the truth. I triangulate.
I ask new leads how they found us and where they were researching, and I actually record the answers instead of letting them evaporate. When someone says "I asked ChatGPT for options and you came up," that is data, even if GA4 never saw it. I watch for the pattern of buyers who arrive already knowing our positioning, already using our language, because that is often the fingerprint of an AI that summarized us well before the human ever clicked. And I look at whether the specific prompts where we improved our presence line up, over time, with movement in the pipeline that matters.
A metric you cannot connect to a decision is not measurement. It is decoration you pay a monthly fee for.
Is that as tidy as a dashboard with a green number? No. Reality rarely is. But I would rather have a rough measure of the right thing than a precise measure of the wrong thing. Precision about a vanity metric is just a more expensive way to fool yourself.
The uncomfortable part nobody says out loud
There is a reason the industry defaulted to citations, sentiment, and referral traffic. They are easy to count. A tool can scrape them, chart them, and sell them back to you every month. Buying stage presence, answer accuracy, and pipeline connection are harder. They require you to actually know your buyer and do the unglamorous work of tying signals to outcomes.
So the industry did what industries do. It measured what was easy and called it what was important.
I understand the temptation. When something new and a little scary shows up, a number, any number, feels like control. But a hammer is not a house. A mention is not a customer. And a rising sentiment score is not a growing business. The whole point of getting recommended by these AI tools is not to be talked about. It is to be chosen.
Track the things that lead to being chosen. Presence when it counts. Position when you are present. Accuracy in how you are described. And an honest, even if imperfect, line from all of that to pipeline. Everything else is weather. Interesting to glance at on your way out the door, and no way to run a company.
The models will keep changing. The tools will keep adding shinier gauges. The question underneath all of it will not change at all. Did being visible make you money, or did it just make you feel seen?


