Most companies right now are asking the same question about AI.
Do we show up when a buyer asks ChatGPT, Claude, Perplexity or any other generative engine for a recommendation? Are we in the answer, or are we invisible?
Nobody can give them a clear answer. There is no playbook.
So I ran a project myself to build one. Not to measure rankings. To measure whether a brand gets recommended by AI to a buyer who never visits its website.
This article gives you the full framework. Every step, every limitation, every honest unknown. If your buyers are talking to AI, this is where you start.
This is long. It is meant to be used, not skimmed.
The Core Position
Prompts are not the new keywords. They are audit instruments.
You cannot observe your buyers inside a generative engine (LLM). You cannot see their chat history. But you can systematically audit what the AI believes about your brand across the problem spaces that matter to your business, and act on what you find.
To measure your visibility in AI generated answers, the honest starting point is this: nobody is handing you clean data. The only tools available today are manual runs across different LLMs, or third party platforms that simulate persona based queries at scale.
1. Strategy
Prompts are fundamentally different from keywords
Keywords are compact. They push a user to condense an entire need into two or three words. Prompts do the opposite. They carry context, constraints, and reasoning. A prompt can be a question, a problem, a comparison, or a full research request.
What makes a prompt genuinely different is that the whole customer journey now lives inside one conversation. Before LLMs, a buyer moved through discovery, research, and comparison across separate channels, often without being fully aware they were doing it. Now all of it can happen in a single thread.
Here is what that looked like in a project I ran for myself, not for any client. I was replacing old decking wood. I dropped two photos into an AI tool and asked which wood type I could use outside that would last and be sustainable.
The AI did not just answer the question. It acted as technical advisor, price comparator, project manager, and safety checker in one conversation. It told me things I had not thought to ask. That one wood type needed stainless steel bolts instead of zinc because of its tannin content. That the planks needed 48 hours to acclimatise before installation. That the frame needed a slight drainage slope even when level.
Problem, research, comparison, decision, and execution planning all happened in one conversation. Only the purchase happened outside it, at a local supplier.
That is the shift. The old discover, research, compare, choose model assumed each stage lived in a different channel. Inside an LLM, all of it can happen in one place, and the transaction still happens elsewhere.
How to categorise prompts
To measure visibility, start by sorting prompts into where they sit in the buyer’s process.
Discovery. Is your brand part of the answer when someone is still exploring. This is the hardest stage to track because the entry point is never predictable. It could be a problem, a research task, a project brief, or simple curiosity. Split it further into problem, research, project, and options.
Solution. Is your brand offered as part of the actual solution the AI proposes. This is easier to track once you know what your solutions are. List your core products, services, and platforms as trackable categories.
Comparison. Do you show up when someone is comparing options. Track prompts structured as “who is better, brand versus brand.”
Selection. This requires input from your sales team on what criteria buyers are actually using today to choose between vendors.
Purchase. Are you present when someone is ready to buy. In almost every case, the transaction itself still happens outside the LLM.
Mapping the prompt set
Before building a prompt set, map the following for each part of your business:
Domain. Which area of the business you are focusing on.
Persona. Who the buyer or influencer is. A persona is rarely one person. It is an organisation with multiple roles inside it, each asking different questions.
Sub-persona. Within that organisation, who specifically is asking. An engineer and a purchasing manager at the same company have fundamentally different questions and need different prompts.
Intent. What the sub-persona is trying to do when they open the conversation. Learn, explore, compare, validate, or decide.
Phase. Where they are in the buying journey at that moment.
Questions. The real questions this sub-persona asks today, collected from the people who hear them daily, not assumed from a desk.
Prompts. Built directly from those real questions, filtered through reach and purchase.
This table cannot be filled in speculatively. It needs a working session with the people closest to the buyer, and it should never be guessed at from the outside.
What makes a prompt worth tracking
The uncomfortable truth is that we cannot yet measure whether visibility in an AI answer leads to a lead or a sale. No tool does that today. Which makes it hard to know when a prompt is representative versus redundant.
The working assumption: being visible is better than not being visible. So prompts earn a place in the set by passing three tests.
It sits at a Reach or Purchase endpoint. It maps to a domain and persona that matter commercially. It represents a class of queries rather than duplicating one already covered.
A prompt fails the set when it sits in the untrackable middle, covers a low value domain, or duplicates the intent of a prompt already tracked.
This rule is provisional. Once real measurement data exists, redundancy becomes empirical. Prompts whose results always move together are measuring the same thing, and one of them can be cut.
2. How To Prioritise
To select a limited set of prompts worth tracking, work through this order. What domain is bringing the most revenue, or is most at risk of losing ground. What personas and sub-personas actually matter to that domain. Are they decision makers inside a real B2B buying journey. What are they already asking your sales team today. Turn those real questions into prompts, weighted toward discovery and purchase. To know whether to go broad or deep, measure what is actually converting through your existing channels first.
One context worth naming for any B2B brand doing this work: a meaningful share of revenue likely flows through distribution partners rather than direct end users. There is a diffuse and largely unmeasured population of end users who may be starting their research inside an LLM long before they ever reach your website or your sales team. These are the buyers you know the least about, and potentially where the biggest visibility gap is hiding.
Breadth versus depth
Breadth means covering many topics with one or two prompts each. Depth means covering fewer topics with many phrasings of the same intent.
Both matter, for different reasons. Depth is what makes the measurement statistically honest. AI answers vary not just across repeated runs of the same prompt but across different phrasings of the same underlying intent. Track one prompt per topic and you cannot tell whether the result reflects your actual visibility or just that one phrasing. Multiple phrasings, each run multiple times, is what separates real signal from noise.
Breadth is what makes the measurement commercially useful. Go deep on one or two topics only and you may measure them perfectly while staying blind to the domains where your biggest gaps actually live.
Start with depth. Fewer topics, several phrasings per intent, each run repeatedly. This calibrates the method. It tells you how much answers vary by phrasing alone, which tells you the real cost of covering any one topic properly. Only once you know that can you responsibly decide how much breadth your budget can afford.
Breadth without depth produces numbers you cannot trust. Depth without breadth produces trustworthy numbers about too little. Resolve it in that order. Learn the required depth first, then scale breadth from there.
3. What You Are Actually Measuring
Choosing prompts is only half the question. The other half is knowing what the AI’s answer actually tells you once you have it.
You are not measuring buyer behaviour. You cannot. You are measuring the AI’s current disposition toward your brand, across four dimensions. This is a genuinely new measurement object. It does not behave like anything from the SEO era.
Reputation. Does the AI have a clear, credible picture of your brand at all. Before an AI recommends anything it needs a coherent understanding of what that thing is. Weak or fragmented signals mean the AI may know you exist but treat you as peripheral rather than credible. This is the foundation everything else sits on.
Perception. What does the AI associate you with. Innovative, reliable, premium, technically deep, sustainable. These associations form from repeated patterns across training sources. They are structural and slow to shift. A buyer asking which vendor is most reliable will get an answer shaped entirely by what the AI has already learned to associate with each name in that category.
Comparison. How does the AI position you against competitors. AI rarely evaluates a brand in isolation. Most real decisions involve a direct comparison prompt, and how the AI frames that comparison shapes the buyer’s shortlist before a single salesperson gets involved.
Recommendation. Does the AI actually recommend you for a specific persona, use case, and context. This is the hardest dimension to earn and the one that matters most. A brand can have strong reputation and clear perception and still lose the recommendation in a specific situation. It depends on persona, region, use case, and decision context. The same brand might be recommended for a large scale build and skipped entirely for a smaller deployment. This is exactly why persona accurate prompts matter so much.
See the AI RepScore for more information — the AIRepScore Framework.
Reach sits at the reputation and perception level. The question is whether you appear at all, and whether you are framed as credible while a buyer is still forming their understanding of the space.
Purchase sits at the comparison and recommendation level. The question is whether you are actively recommended once a buyer is close to a decision.
The two knowledge layers
When an LLM answers a question about your brand, it draws on two different sources, and they behave very differently.
Trained knowledge is what the model learned during training, from websites, articles, product pages, reviews, and technical documentation. It is frozen at a fixed cutoff and only shifts when a new model version releases, typically every six to eighteen months. A gap here is slow to close. It needs long term content strategy and consistent presence across many external sources.
Live retrieval is what happens when a model searches the web at the moment of the query and pulls in current content. This is where fresh web signals matter, and gaps here respond faster to updated documentation and active publishing.
When you find a visibility gap, work out which layer it lives in. Absent from trained knowledge means a long term strategy problem. Absent from live retrieval means a faster fix. Running the same prompt with and without web retrieval active will tell you which one you are looking at.
4. A Real Prompt Set From A Real Journey
Instead of inventing a hypothetical, let me show you what these prompt categories look like using a journey I documented end to end: my own.
Earlier I mentioned replacing the decking at my home in the Netherlands. That project became an accidental case study. It began with two photos and one open ended request, “I want to replace this,” and it ended weeks later with a completed physical deck, a documented materials and tools list, and purchases across multiple retail brands. Every brand touchpoint, every comparison, every purchase decision, and every mid-project correction happened inside one ongoing AI conversation.
Here is how the prompts from that single journey map onto the framework.
Reach: Discovery
I want to replace these woods, which wood type can I use for outside that will last long and is sustainable? (with two photos attached)
What are the differences between hardwood and composite decking?
What do I need to check on the existing frame before I put new planks on it?
How much does a decking replacement like this typically cost?
Reach: Solution
What materials and tools do I need for this project, make me a complete list.
Where can I buy this type of hardwood decking in the Netherlands?
Which fasteners work with this wood type?
Purchase: Comparison
Compare the prices and quality of decking planks across these retailers.
Which of these suppliers is the better option for the amount I need?
Purchase: Selection and validation
Is this wood the right choice given my situation and budget?
Are these the right bolts for this wood?
Purchase: Decision support
Make me the final shopping list with quantities.
Give me the step by step plan for the installation.
This is not going as planned, what do I do now? (mid-project corrections, several times)
Three things in this journey should concern every brand.
First, no retailer’s name appeared until the comparison stage, and even then, the AI introduced the names, not me. I did not know which suppliers to consider. The AI formed my shortlist for me.
Second, the AI added expertise I did not know to ask for. It told me the wood’s tannin content required stainless steel bolts instead of zinc. That the planks needed 48 hours to acclimatise before installation. That the frame needed a slight drainage slope even when level. Every one of those answers shaped what I bought and where.
Third, the retailers who won and lost my money had zero visibility into any of it. From their side, I appeared as a customer who walked in already decided. The entire decision happened somewhere they could not see.
This is a B2C project, and the framework in this article is built for B2B. But that is exactly the point. If a complete purchase journey can collapse into one AI conversation for a deck, it is already happening for software, infrastructure, and services, where buyers do far more research before they ever contact a vendor.
A real tracked set for your business follows this same structure, thirty to fifty prompts, built from real buyer questions, with every intent covered by several phrasings, following the depth principle above.
5. Measuring It
At this early stage, the honest bar for visibility is simple. Your name appearing in an AI generated answer counts. As measurement matures, move toward the four dimension model above, tracking not just whether you appear but what the AI associates you with, how it positions you against competitors, and whether it actively recommends you for the contexts that matter.
For discovery prompts, track inclusion. Are you named in the answer at all.
For purchase prompts, track recommendation. Are you specifically suggested as an option for that buyer’s context.
As the method matures, add accuracy of description, sentiment, whether your own content is cited or only third party descriptions of you, and whether the answer drew on trained knowledge or live retrieval.
Turning this into something usable
Since the purchase itself does not happen inside an LLM, you need to track what happens on your own platforms once someone arrives from one. Add a simple field to your existing forms asking whether and how someone found you through an AI tool. That single field creates a first party signal connecting AI activity to what actually happens on your site.
Run an actual customer survey alongside it. No tool replaces asking real buyers how and when they are using AI in their own process.
Over time, look for correlation between visibility and business outcomes. Not direct attribution, that is not possible yet. Directional correlation is enough to tell you whether the work is moving the right numbers.
What the reporting should look like
Monthly: a visibility score per domain, the percentage of reach prompts where you are included and the percentage of purchase prompts where you are recommended, tracked over time. A competitive displacement view, showing which competitors appear in the answers where you are absent. A gap list of the specific prompt categories producing weak or zero visibility, which becomes the direct input for content and PR decisions.
Quarterly: review all of it against the first party signal from your forms and channel data, looking for movement that lines up.
6. Running An Internal Research Project
If you work inside a large organisation, you likely have the scale to run this as a proper internal experiment across teams, locations, and time zones.
Name the limitation honestly. Employees are not the persona. Someone who already works for the company will shape their prompts around internal knowledge and internal language. Their results tell you what the AI says when the prompt is run. They do not tell you how a real buyer with zero prior knowledge experiences that same answer.
An internal project is useful for exactly one thing: building a baseline audit of what the AI says, at scale, across platforms, geographies, and time.
To get genuine buyer signal, you need one of three things instead: structured research with real buyers, a third party tool that simulates persona based queries, or carefully constructed synthetic persona prompts validated against real buyer language from your sales team. All three cost more than an internal project. The internal project remains the right place to start, as long as everyone involved knows what it can and cannot tell you.
On sample size. Ten runs per prompt offers a reliable balance of statistical validity and practical effort. More than that does not meaningfully shift results. But the prompts themselves must vary. Running one prompt ten times tells you about that phrasing, not about a domain. You need several prompts representing the same intent, each run ten times.
The minimum structure for statistical validity sits around 45 participants. Aim higher for geographic diversity and a buffer for drop off, somewhere in the range of 150.
A workable experiment structure: thirty prompts across reach and purchase categories, each run ten times per platform, across five major AI platforms, with 150 participants spread across at least three geographies, each running the full set on one assigned platform. That produces well over four thousand recorded answers in a single round, a dataset most external tools cannot replicate at that cost.
Participants run the exact prompts they are given, at different times of day and in different locations, and record the exact answer verbatim. No summarising, no interpretation. Someone owns the project end to end: defining the prompts, designing the recording template, recruiting participants, and consolidating the results.
7. Who Should Own This
AI visibility does not belong to one team by default, but someone needs to own the decision of what to track and how to read it. In practice this sits closest to whoever already owns brand strategy. They decide which prompts to track, which domains to prioritise, and what the numbers actually mean for the business. Domain, persona, and sales input should feed that decision, but the decision itself needs one clear owner.
Acting on what the measurement shows is a wider conversation involving marketing, SEO, PR, and content. That comes after the measurement is consistent and trusted, not before.
A workable cadence: monthly runs across priority platforms with results logged and gaps identified. Quarterly review of what changed, which gaps closed, and what moved. Ongoing correlation between visibility trends and the business outcomes you can actually measure.
Start with one domain. Prove the operating model works on something small before scaling it across the business.
8. What This Does Not Replace
This is a new layer, not a replacement for anything that already works.
It does not replace SEO. Organic search still drives real traffic and real conversion, and the two systems serve different channels and different moments in the buyer’s journey.
It does not replace brand tracking. Traditional brand health measurement still captures human perception in ways AI measurement cannot reach.
It does not replace your sales pipeline. The best signal of what buyers are actually thinking is still the conversations happening between buyers and your sales team every day. That data should feed this framework, not be replaced by it.
It does not replace first party data. What happens on your own site and platforms remains the most direct signal of buyer behaviour you have. AI visibility measurement sits upstream of that. It tells you what might be shaping a buyer before they ever arrive.
What it does uniquely is capture the part of the journey that used to be invisible entirely. The conversations happening inside AI tools before a buyer ever visits a website, fills in a form, or speaks to a person. Nothing else currently reaches that.
9. What This Framework Cannot Tell You Yet
Worth saying plainly, because most work in this space overclaims.
You cannot yet track what AI visibility directly leads to. No tool does that today. You cannot confirm which prompts are representative versus redundant until real measurement data exists. The domain, persona, and question mapping cannot be filled in from a desk, it needs a working session with the people closest to real buyers. Any illustrative prompt set is a draft until it is built from real buyer language. An internal employee project tells you what the AI says, not how a real buyer experiences it.

