How do you measure AI visibility? Moving from keywords to decision questions

How to measure AI visibility: Moving from keywords to decision questions

We explain why keyword rankings can't measure AI visibility, which indicators a decision-question-based measurement uses, and how Recro's four-pillar simulation works.

AI visibility is measured by how often, in what position, on what grounds and based on which sources a brand is recommended in the decision questions its target buyers ask AI. The unit of measurement is the decision question, not the keyword. Because answers change with every run, a single rank number means little; what matters is the pattern of appearance rate, framing and sources across repeated runs. Accurate measurement starts with brand-specific personas and brand-specific questions.

An industrial automation company’s marketing report is green: most target keywords rank on Google’s first page. In the same period, the sales team notices fewer new inquiries, and the customers who do arrive open with “we compared you with competitor X.” Nobody sees this question in the report: when a plant manager asks AI “Which automation integrators are compatible with our existing PLC infrastructure and have a service network in Turkey?”, who is in the answer?

This article explains how to measure that: with which unit, which indicators, how many repeated runs and which pitfalls to avoid. At the end, you’ll see why Recro builds this measurement on four brand-specific pillars.

Why don’t keyword rankings show AI visibility?

A ranking report measures position on a search engine results page. An AI answer has no results page. It has a few brands, a rationale and a few source links. What the buyer sees isn’t a list; it’s a recommendation.

The shift in buying behavior turns this difference into a commercial problem:

  • According to Forrester’s 2024 Buyers’ Journey Survey, 89% of B2B buyers have adopted generative AI and name it one of their top sources of self-guided information in every phase of the buying process.
  • In G2 research published in April 2026, 51% of B2B software buyers say they now start research with an AI chatbot more often than with Google, up from 29% in April 2025. 69% say they chose a different vendor than originally planned based on AI chatbot guidance.
  • 6sense’s 2024 Buyer Experience Report finds that 81% of buyers already have a preferred vendor at first contact, which usually comes after roughly 70% of the journey is complete.
  • Gartner reports that B2B buyers spend only 17% of their total buying time meeting potential suppliers, and its March 2026 survey found that 67% of buyers prefer a rep-free buying experience.

The shortlist forms in a space the sales team can’t see, and a growing share of that space is made of AI answers. Ranking reports don’t measure it.

There’s a technical reason too. Google has said AI Mode uses a “query fan-out” technique that breaks a single question into many sub-queries. The sentence the user types becomes dozens of background sub-questions the brand never tracks. A keyword list doesn’t cover them.

The decision question: the new unit of measurement

A decision question is a question that directly affects a buyer’s decision to put a supplier on the shortlist or take it off. It has four components: who asks, what triggers the question, the condition the answer must meet and the acceptance threshold.

Keyword vs decision question
Keyword: “automation integrator”
Decision question: “Which integrators can modernize our line without replacing our Siemens S7 infrastructure, offer 24-hour service in Izmir and have food industry references?”

The second sentence contains a persona (plant manager), a trigger (modernization), conditions (infrastructure compatibility, region, sector references) and an implicit threshold (24-hour service). AI builds the answer around these components. A brand’s visibility is won or lost separately in each of them.

We explain how decision questions are collected and mapped in what is a decision question. The critical point for measurement: unless the question set is written in the buyer’s language rather than the brand’s, the measurement only tests the questions the brand asks itself.

Which indicators are measured?

AI visibility can’t be reduced to a single number. A measurement that supports decisions reads at least six indicators together.

Indicator Question Why it matters
Appearance rate In what share of answers across repeated runs does the brand appear? Far more stable than position in a single answer
Position band Is the brand the first recommendation, mid-list or among “others”? Read as a band, not an exact rank
Framing In what role is the brand mentioned: preferred, specialist, alternative, budget option? Two brands mentioned equally often can be perceived very differently
Source Which sources does the answer rely on when mentioning the brand? Shows where action should go
Accuracy Is what’s said about the brand current and correct? Misrepresentation can cost more than invisibility
Persona coverage In which buying committee roles’ questions does the brand appear? Being strong with one persona doesn’t win the decision

We covered how these indicators are defined as measurement criteria in AI visibility measurement criteria. The focus here is the unit they’re measured on and how they’re read together.

Why is a single run misleading?

In research published by SparkToro and Gumshoe.ai in January 2026, hundreds of volunteers ran the same 12 prompts through ChatGPT, Claude and Google’s AI answers nearly 3,000 times. The result: the chance of two answers giving the same list of brands was below 1 in 100. The chance of the same list in the same order was closer to 1 in 1,000.

This says two things. First, a report like “we’re number 3 in AI” is meaningless. Second, what varies is the order; what’s relatively stable is whether the brand enters the answer across many runs. That’s why measurement is built on appearance rate, position is read as a band, and results are tracked as a time series rather than a snapshot.

Which AI tools should be measured?

Each tool relies on different sources. Studies of AI answer citations show that only a small share of the domains cited by ChatGPT and Perplexity overlap. Measuring in a single tool mistakes that tool’s source preferences for the brand’s overall position.

A practical set for B2B: ChatGPT (with and without search), Google’s AI Overviews and AI Mode, Gemini, Perplexity and, where common in corporate environments, Microsoft Copilot. Weighting depends on which tool the target buyer actually uses. A corporate procurement team and a software engineering team don’t use the same tools.

What do AI visibility tools on the market measure?

Platforms such as Semrush’s AI visibility modules, Ahrefs Brand Radar, Profound and Peec AI track how often a brand is mentioned across defined prompt sets, which sources are cited and how the brand compares with competitors. They’re valuable for broad category tracking and trend monitoring.

Their limit lies in where the measured questions come from. These tools mostly work with category-level prompts that are the same for everyone, or prompts users write themselves. In B2B, the question that decides the deal is often specific to one brand, one sector and one decision-maker. Appearing for “best CRM software” doesn’t mean appearing for “a CRM for a pharmaceutical distributor with a 200-person field sales team running on SAP.” The second question isn’t on any tracking tool’s ready-made list; the persona and the question have to be built first.

Measurement pitfalls

  • Measuring the brand’s own language. If marketing writes the question set, questions contain the brand’s category names and results look better than reality.
  • A single persona. Being strong with one member of the buying group hides disappearing from the vetoing member’s question. According to 6sense, the average B2B buying group has 11 people.
  • A single language. An exporting brand’s position in English questions can differ sharply from Turkish ones. We showed this in Turkish vs English prompts.
  • Confusing the number with the cause. If appearance rate is low, measurement doesn’t stop there. The real value is finding why the brand is missing: missing information, wrong framing or no external confirmation?
  • Reading without verification. AI output isn’t evidence on its own. Claims in answers should be checked against their counterparts on the web before a report is written.

How does Recro run this measurement?

The Recro Insight Model builds this measurement as a brand-specific simulation. Its value doesn’t come from the AI tools used; it comes from whose eyes the tools are approached through, which questions are asked and what the answers are read against.

One simulation · 4 key pillars · 90+ steps

What sets Recro apart is four pillars built from scratch for each brand. Together they form a single simulation of 90+ steps, with each pillar producing the input for the next.

  1. Persona builder: Decision-maker profiles that represent the brand’s real buyers, specific to its sector and sales structure.
  2. Question builder: The brand-specific decision questions these people ask AI along their buying journey.
  3. Report builder: A report that reads the answers against the brand’s goals and competitors, together with the sources behind them.
  4. Action recommendation builder: A prioritized list of actions that closes the sector-specific signal gaps.

What each pillar contributes to measurement:

  • The persona builder defines on whose behalf the measurement is made. Personas are decision-maker profiles derived from the brand’s real buyer structure, not generic titles, and can optionally connect to the brand’s anonymized CRM data. Single persona–many questions, many personas–single question and many personas–many questions setups cover the whole buying group.
  • The question builder derives questions from persona attributes and sequences them in a natural buying journey flow: from discovery to screening, from verification to risk questions. The measured questions become the buyer’s sentences, not the brand’s.
  • The report builder reads the answers with the indicators above: answer analysis, the order in which the brand enters the answer, the links that source the conversation, estimated visibility probability and a KPI assessment set against the brand’s own goals.
  • The action recommendation builder turns findings into a sector-specific list of LLM signals and a prioritized action plan. It also determines which equivalent content to produce against leading competitor content.

Measurement isn’t left as a one-off snapshot. Daily outputs become layers of data, findings, insight and strategic insight over time. In the monthly reality check, the web counterparts of prominent answers are verified and the report is finalized through human review. Reasons for not appearing are classified as Communication, Content or Authority GAPs.

Sample finding format · illustrative

  • Persona: Plant manager · Food industry · Turkish
  • Question: Screening · Modernization compatible with existing infrastructure
  • Appearance: Low; two competitors are in the top band in most runs.
  • Source: Competitors’ compatibility lists and case pages are cited.
  • Class: Content GAP
  • Action: Publish supported PLC families, regional service coverage and food industry cases as text.

Executive summary

  • The unit of AI visibility measurement is the decision question. Keyword rankings don’t show the space where shortlists form.
  • Because answers change with every run, appearance rate, framing, sources and persona coverage are tracked together over time, not rank.
  • The value of measurement comes from brand-specific persona and question design, not the tool. Alongside the number, the reason for invisibility and the action should be reported.

Frequently asked questions

Is AI visibility the same as SEO visibility?

No. SEO visibility measures position on a results page. AI visibility measures whether, and on what grounds, the brand is recommended in the answer to a buyer’s decision question. They’re related, but one doesn’t replace the other.

How many times should a question be run?

A single run isn’t enough. Because answer lists rarely repeat, each question is run several times in each tool and the result is read as a rate. The number depends on the size of the question set and the confidence level required.

Is there an answer to “what’s our rank in AI?”

An exact rank isn’t reliable. What matters is the rate and position band in which the brand appears, the role it’s given and how that changes over time.

Are off-the-shelf AI visibility tools enough?

They’re useful for category-level trend tracking. Measuring the brand-, sector- and decision-maker-specific questions that decide B2B deals requires building those personas and question sets first.

How often should measurement be repeated?

Daily monitoring catches sudden shifts, monthly reviews clarify priorities and quarterly reading supports strategic decisions. Interim measurements should follow model updates and major content changes.

When your critical customer asks AI, is your brand in the answer? Recro builds a sample simulation with personas and decision questions specific to your brand and shares the result in a demo report.

Request a demo insight report →

Sources


This article was prepared with AI assistance and published under the review of Mehmet Semih İpek.