How to measure AI search visibility
Your buyers ask ChatGPT or Perplexity for a shortlist before they search, and your analytics only see the clicks, never the answers that left you out. You can measure it, but not with a screenshot. This is the method: which questions to ask, how many times, which numbers to keep, how to tell a real change from noise, and what to do with the result. It works with a spreadsheet and patience, or with an agent doing the asking for you.
Direct answer
How do you measure AI search visibility?
Ask the AI engines your buyers use a fixed set of buyer questions, several times each, on a regular schedule, and count. Report the share of answers that mention you (mention rate), your share of all brand mentions against named competitors (share of voice), and the domains cited in answers that left you out. Call a change real only when it is larger than normal run-to-run variation.
The numbers worth keeping
Six numbers per engine, per run. Every one is computed from the same stored answers, so you can recompute history when you fix an alias.
| Number | How to compute it | What it tells you |
|---|---|---|
| Mention rate | Answers that name you, divided by all answers | How often you are in the conversation at all |
| Share of voice | Your mentions, divided by mentions of every tracked brand | Your slice of the shortlist against named competitors |
| Rank | Your position among tracked brands, by first mention | Whether you are the first name or an afterthought |
| Citation rate | Answers that link to one of your domains | Whether the engine uses your pages as a source |
| Cited where you are absent | Domains cited in answers that did not mention you | Where the engines get their shortlist, and where you are missing |
| Prompts gained and lost | Prompts where you went from zero mentions to some, or back | Which buyer questions moved, so you know where to read |
Why one spot check tells you nothing
The usual first attempt is to type your category into ChatGPT, see your name, and feel fine. Then a colleague asks the same thing an hour later and gets a different list. Generated answers are sampled, so the same question can produce different brands on different runs, and each engine retrieves and ranks sources its own way. The consumer apps can also personalize: ChatGPT, for example, can use saved memories and past chats when you turn that on. A single answer is an anecdote about one run on one account.
- Measure rates, not yes or no: "mentioned in 7 of 9 answers" survives a rerun, "it mentioned us" does not.
- Measure every engine your buyers use separately. Visibility in one says little about another.
- Use a clean, repeatable setup rather than your own logged-in account, so last week and this week are comparable.
Step 1: write the prompt set
The prompt set is your keyword list and your time series at once, so it deserves more care than anything else here. Write 10 to 30 questions the way a real buyer types them before choosing, not the way your marketing team would phrase them. Give each one a permanent id. When you reword a prompt, it becomes a new prompt with a new id, because a reworded question is a different measurement and quietly swapping it breaks every trend line that ran through it.
- Category questions, where you have to earn the mention: "best invoicing tool for freelancers".
- Comparison and alternatives questions: "X vs Y", "alternatives to a competitor".
- Problem questions that lead to your category: "how do I stop chasing late invoices".
- A few branded questions, where the test is accuracy rather than presence: "does X have an API".
- Weight the set toward category questions. A prompt that names you will mention you, which proves nothing.
Step 2: name the field and choose the sampling
List your own brand with every alias and domain it goes by, then the three to eight competitors a buyer would actually weigh you against. Count mentions for all of them, because your number alone has no context: a mention rate of 30 percent is weak if two rivals sit at 80 and strong if nobody else clears 20. Then pick the sampling. Ask each prompt at least three times per engine per run, with web search turned on so the engine answers from what it retrieves today rather than only from training data. If you query through an API rather than the consumer app, say so in every report: the API gives you a consistent panel for trends and sources, not a copy of a buyer's screen, and API and app answers can name quite different brands.
- Calls per run are prompts times engines times samples: 20 prompts, 5 engines and 3 samples is 300 answers.
- At 3 samples, 20 prompts gives you 60 answers per engine, enough to compute a rate you can compare.
- Aliases matter. A brand that is also a common word needs a stricter match or it will be over-counted.
- Store every raw answer and every cited link, not only the tally, so you can recount after fixing an alias.
Sources: Web search for models, OpenRouter docs · Scraped AI answers vs API results, Surfer (2026)
Step 3: separate a real change from noise
This is the step most reports skip, and it is where most false wins come from. With 60 answers per engine, a move from 12 mentions to 18 looks like 50 percent growth, but a standard two-proportion test puts it at about 1.3 standard errors: well within the variation you get by rerunning the same week. A move from 15 to 27 is about 2.3 standard errors, which is worth calling a trend. A simple rule holds up: treat a change of two standard errors or more as likely real, and label everything smaller as noise, however good the story sounds.
- Keep the series comparable. Changing the model behind an engine, the sample count or a prompt breaks comparison; note the date and mark the break in the next report.
- Read the answers behind every change before reporting it. Counting misses paraphrases and catches false positives.
- Check how you were described, not only whether you were named. A mention with the wrong price or a dead feature is worse than silence on a money prompt.
- Keep your own change log: every page you shipped or listing you claimed, with its date, so you can line it up against the numbers.
Step 4: act on the citations
The most useful output is not your mention rate. It is the list of domains the engines cite in answers that left you out. Those pages are where the shortlist comes from: a review site category, a comparison article, a community thread, a documentation page the engine keeps misquoting. Each one points at a concrete move. Keep recommendations few and specific, tie each one to the prompt, the engine and the evidence, and label a guess as a guess. The research paper that coined the term generative engine optimization found that some content changes raised visibility in its benchmark, which is encouraging, but there is no switch that guarantees a mention.
- A review or directory site keeps getting cited: get listed there, with a complete and accurate profile.
- A comparison page wins a prompt you lose: publish your own plain, honest comparison for that question.
- An engine states a wrong fact about you: answer the question directly on your own site, where it can be retrieved.
- No more than three recommendations a week, then watch the prompts they targeted and report what happened, including when nothing moved.
Sources: GEO: Generative Engine Optimization (Aggarwal et al., arXiv)
Doing it every week without doing the counting
None of the method is hard. The problem is the repetition: five engines, twenty prompts, three times each, every week, with the tally, the noise test and the reading done before Monday. That is the homework, and it is exactly the kind people drop after the third week. Qoren's AI visibility tracker template runs this method as an always-on agent: it collects on a schedule, stores every answer and citation, flags which changes are likely real, and sends a one-page brief with at most three recommendations. The fixing is still yours.
Running the method for clients
Agencies get asked this question by their clients before anyone else does. The method does not change when you run it for someone else, but the setup should: keep one prompt set, one competitor list and one tracker per client, each in that client's own environment with its own keys and budget, so no client's answers or history mix with another's. The Monday brief comes to you, you read it, and you forward it under your name with your own recommendations on top.
Related guides
AI visibility for agencies
AI visibility tracking for agencies: a weekly GEO brief for every client
Offer AI visibility tracking to every client: one tracker agent per client environment, weekly mention rates across AI engines, and a brief you forward.
Read guideFor agencies and resellers
You sell the agents. We do the homework.
A platform for agencies to build, deploy and resell managed AI agents to SMB clients: an isolated environment per client, one console, your own invoice.
Read guideMoney
AI agent spend control
How AI agents run up unexpected bills: retry loops, growing context, oversized tool payloads. Why alerts arrive too late, and what a real hard stop looks like.
Read guideTools, templates and next steps
AI search visibility
Track whether ChatGPT, Claude, Perplexity, Grok and Gemini mention your brand when buyers ask, how that changes each week, and which sources they cite.
OpenAI visibility tracker template
OpenTrack Brand Mentions in ChatGPT: A New Qoren Template
Our new agent template asks five AI engines your buyers' questions every week, counts who gets mentioned, and flags which changes are real. How it works.
OpenFrequently asked questions
How many prompts do I need to measure AI visibility?
Between 10 and 30 is the useful range. Fewer and one odd prompt swings the whole number; more and the cost and the reading grow faster than the insight. Pick the questions that carry commercial weight, weighted toward category questions that do not name you.
How often should I measure?
Weekly suits most businesses. Visibility tends to move over weeks rather than hours, and a weekly run keeps the cost modest while still catching a change within days of the page or listing that caused it. Review the prompt set and the models behind each engine about once a month.
Is API data the same as what people see in the ChatGPT app?
No, and the difference can be large. An API call with web search on is consistent and repeatable, which makes it good for trends. The consumer apps can add memory, personalization, their own instructions and their own search behavior, and one published comparison found API and app answers named the same brands only about 16 to 24 percent of the time. Treat the numbers as a trend line on a fixed panel, not a screenshot of one buyer's screen.
What is a good mention rate?
There is no universal benchmark, because it depends entirely on your prompt set and your field. Compare yourself with your own history and with your named competitors on the same prompts. Share of voice against competitors is usually the more honest number to put in front of a team.
Can an agency measure AI visibility for several clients?
Yes. Run the same method once per client: a separate prompt set, competitor list and tracker for each, kept in that client's own environment so histories never mix. You receive each weekly brief and forward it under your name. The agency guide covers the setup, the pricing and what the client actually sees.
Can I do this with a spreadsheet?
Yes. Paste each answer into a sheet, mark who was named and which links were cited, and compute the rates. It works, it is slow, and it is the part people stop doing after a few weeks, which is why the storing and counting are worth automating even if you keep the reading.
Let an agent do the asking and the counting.
The AI visibility tracker template runs this method every week, stores every answer and citation, and sends a one-page brief with at most three recommendations.