AI Search Visibility: How to Measure It and Steer It

Dominik Breitbach founded taismo, an SEO and GEO agency from Munich, and works as its Lead SEO Strategist. He builds the measurement panels that tell clients whether AI systems name them, and which competitor gets named instead.
🔄 Last updated: 15 August 2026
AI search visibility is the share of your customers’ questions where an AI system names your company or cites one of your pages. It has no position one and no ranking report behind it, so it gets measured as a rate: You run a fixed set of prompts on a schedule, record who gets named and who gets linked, and read the change across months. Three data sources back that panel up: Generative AI impressions in Search Console, referral traffic from AI hosts, and AI crawler hits in your server logs.
👉 Want a number instead of a hunch? We build the prompt panel and the baseline as part of a GEO focused SEO audit.
Mention, citation and click are three separate metrics
Most reporting on AI search visibility mixes three things that behave differently. Separate them before you measure anything.
A mention puts your brand name in the answer text without a link. A citation lists one of your URLs as a source. A click happens when a reader follows that link. Each layer is smaller than the one above it, and each one is visible to a different tool.
The gap matters in practice. A model can recommend your company in three sentences and cite a directory page instead of your site. Your analytics sees nothing, your Search Console sees nothing, and the buyer still arrives two weeks later typing your brand name into Google. Reporting clicks alone therefore understates AI visibility, usually by a lot.
So the mention is your primary metric, the citation your secondary, and the click the outcome you check last. That order inverts classic search reporting. Generative engine optimization describes the work that produces those mentions; this guide covers how you count them.
Four data sources, and the question each one answers
AI search visibility has no single source of truth. Four sources exist, they overlap, and each is blind in a different place. This is the table we hand clients when a GEO engagement starts.
| Data source | Question it answers | What it cannot tell you |
|---|---|---|
| Prompt panel | Do AI systems name you, cite you, and who do they name instead? | Nothing about real demand. A prompt is not a keyword with a search volume. |
| Search Console | How often were links to your site shown inside AI Overviews and AI Mode? | Clicks, click through rate, position, and every system that is not Google. |
| Referral data | How many people actually arrived on your site from an AI host? | Mentions without a link, and the prompt that triggered the visit. |
| Server logs | Which AI crawlers request which pages, how often, and with what status code? | Whether the fetch ever turned into an answer. |
Run all four and the picture closes. Run one and you will end up reporting whichever metric happens to move.
Search Console now reports generative AI impressions
Google changed the measurement situation in June 2026 with the Search Generative AI performance report in Google Search Console. It shows how often links to your site were shown inside a generative AI feature, broken out from the rest of your search data for the first time.
Four limits define what that report can do:
- Impressions only. No clicks, no click through rate, no average position.
- Two surfaces. AI Overviews and AI Mode. Search Labs experiments are excluded, and Discover has its own separate report.
- Four dimensions. Page, country, date and device, with canonical URLs as the page unit.
- Partial rollout. Google states that not all properties have access yet, that sites need enough impressions for data to show, and that sites which opted out of generative AI features are excluded.
In the regular Performance report, AI features stay inside the Web search type, which is why your normal totals already contain them. Two counting rules change how you read those totals. An AI Overview occupies a single position, and every link inside it is assigned that same position, so a strong AI citation on a page two topic can quietly drag your average position around. And an impression only counts once the link has been scrolled or expanded into view, which means a collapsed AI Overview generates nothing at all.
AI Mode counts differently again. Google treats a follow-up question inside AI Mode as a new query, so its impressions, position and clicks are attributed to that new query. Long conversations therefore inflate query counts and flatten your click through rate compared with a classic result page.
Four gaps matter more than the impression count itself. There is no query dimension, so the report names the page that was shown and never the question that produced it. Pages are grouped by canonical URL, so variants roll up and stop existing as separate lines. The table inherits the 1,000 row limit of the regular Performance report, which truncates large sites quietly. And the dates run on Pacific Time with the newest days flagged as preliminary, so a European reporting month never lines up cleanly at its edges and chart totals can differ from table totals.
Two readings are therefore off limits: A click through rate, because the report holds no clicks, and the attribution of any single visit to an AI surface. It is a coverage signal, and it is the only one Google gives you.
Build a prompt panel that survives a non-deterministic system
Ask the same AI system the same question twice and you can get two different answers with two different sources. That is how the systems work, not a setting you can switch off, and it is why a single check proves nothing. A panel handles it through repetition.
Four rules make a panel usable:
- Write 20 to 40 prompts in your customers’ words. Pull them from sales calls, support tickets and your own keyword research. Mix four types: Category questions (“best provider for X”), problem questions, comparison questions and brand questions.
- Run each prompt three times per system per cycle. Three systems, 30 prompts, three repetitions is 270 recorded answers per month. That is roughly two hours of manual work, or one scheduled tool run.
- Start every run in a fresh session. Memory and personalization off, location consistent. A logged in account that already knows your brand will name you and tell you nothing.
- Record verbatim. The answer text, every brand named, every URL cited, the system and the date. The raw log is what makes a disputed number checkable six months later.
Panel size decides how much noise you are reporting. With 10 prompts, one answer flipping moves your rate by 10 percentage points. With 40 prompts, the same flip moves it by 2.5. Below 20 prompts you are reporting weather rather than climate.
One more discipline: Freeze the panel. Adding prompts because the current ones look bad destroys the comparison you spent months building. New prompts belong in a second panel with its own baseline.
The five numbers worth putting in a report
Five numbers cover AI search visibility without inventing a score.
- Presence rate. Prompts where you are named, divided by prompts run. Your headline number.
- Citation rate. Prompts where one of your URLs is linked, divided by prompts run. Always lower than presence rate, and the gap between the two is itself a finding.
- Share of voice. Your mentions divided by all brand mentions across the panel. This is the number that survives a growing market.
- Competitor set. The ranked list of brands named instead of you.
- Cited pages. Which of your URLs the systems actually pick. Usually fewer than you expect, and rarely your homepage.
Two of those five are work orders rather than reports. The competitor set shows who owns the citation slot today, and the cited pages show which format your domain gets trusted for. In the panels we run, the cited pages are almost always the ones that answer a single question completely, not the ones that cover a broad topic.
Tools will offer you one combined visibility score instead. Report the underlying rates as well. A score that folds mentions, citations and sentiment into one figure cannot be reconstructed, compared against another vendor’s score, or explained to a management team that asks how it was calculated.
Sentiment is the sixth number most vendors add. Record it as verbatim quotes rather than as a score: The finding you can work with is the phrase a model uses about you, something like “solid but expensive”, not a value between zero and one from a classifier you cannot inspect. Turn it into a rate only once your presence rate is high enough that there is something to be positive about.
Work the five numbers on a real panel
Definitions get argued about; arithmetic does not. Take the panel described above: 30 prompts, three systems, three repetitions, which is 270 recorded answers per cycle. Worked through with example figures, a first month reads like this.
| System | Answers run | Named you | Presence rate | Cited one of your URLs | Citation rate |
|---|---|---|---|---|---|
| ChatGPT | 90 | 31 | 34% | 12 | 13% |
| Google AI Mode | 90 | 44 | 49% | 27 | 30% |
| Perplexity | 90 | 38 | 42% | 24 | 27% |
| Panel total | 270 | 113 | 42% | 63 | 23% |
Read the gap before the level. The distance between the two rates is widest on ChatGPT, where the model names you from what it already learned rather than from a page it just fetched. That points at different work than a low presence rate does: Your name travels, your pages do not.
Share of voice runs on the same log with a different denominator: Your mentions divided by every brand mention in the panel. If those 270 answers contain 690 brand mentions and 113 of them are yours, your share of voice is 16 percent. A single counting rule decides that figure, so fix it before the first run: Count a brand once per answer, not once per occurrence. Otherwise a competitor named four times in one paragraph outweighs a company named once in four separate answers.
Pool by counting, never by averaging the rates. The three system rates of 34, 49 and 42 average to 42 here only because every system ran the same 90 answers. As soon as one run fails or one system gets an extra repetition, the average of rates drifts and the pooled count does not.
When a change is a change and not noise
A presence rate is a proportion, and proportions from small samples wander on their own. The standard error is the square root of p times one minus p, divided by n. With p at 0.42 and n at 30 prompts that comes to 0.09, so the 95 percent band around your 42 percent runs from roughly 24 to 60 percent.
Two consequences follow. A move from 42 to 47 percent on a 30-prompt panel sits inside the noise and carries no information on its own. And halving the band costs four times the panel, because n sits under a square root: 120 prompts bring the standard error down to 0.045. For most companies the better trade is to keep 30 prompts and judge the direction across three cycles, which is the same sample size logic that governs any A/B test.
Two rules keep the arithmetic honest. First, the prompt is your unit, not the answer, because 270 answers are not 270 independent observations: Three repetitions of one prompt are correlated by design. Score the prompt from its repetitions, so named in two of three runs counts as 0.67, and let the 30 prompts be your n. Second, name the prompts won and lost next to the percentage. “We gained four prompts, two of them comparison questions” survives a follow-up question. “Up five points” does not.
We build the prompt panel, run the baseline across ChatGPT, Google AI Mode, Perplexity and Copilot, and turn the result into a work order. taismo is an SEO and GEO agency from Munich, ranked fifth in the German SEO Contest 2026 and second among Munich online marketing agencies by Agenturtipp.
Read AI referral traffic without over-reading it
Clicks from AI systems arrive with a referrer you can see. Hosts such as chatgpt.com, perplexity.ai and copilot.microsoft.com show up as referral sources in any analytics tool. Google’s AI surfaces do not: Clicks from AI Overviews and AI Mode arrive as ordinary Google organic traffic, so your referral list will always understate the total.
The volume is small, and that is normal. Google reports that clicks from result pages with AI Overviews are higher quality, with users more likely to spend more time on the site. Judge this traffic on conversions and engagement, not on session count: Fifty sessions a month from AI hosts can outproduce a thousand from a weak channel.
The second reading is indirect and more useful. When a model names you without linking you, the buyer searches your brand name afterwards. Branded impressions in Search Console are therefore the closest proxy you have for the mention layer. Watch that line against your panel’s presence rate: When both rise in the same quarter, the mentions are reaching people.
Server logs answer the question that comes before visibility
A page that no AI crawler ever fetched cannot be cited. Server logs are the only place that fact is recorded. Every line carries the user agent, the requested URL, the status code and the timestamp, which is enough to answer three questions.
- Which pages do AI crawlers request? Compare that list against the pages you actually want cited.
- How often, and did the rhythm change? A new page fetched within days behaves differently from one ignored for a month.
- What do they get back? A 200 with complete HTML, or a 403 from a security layer.
The third question catches more sites than the first two. Your robots.txt is visible in your repository, but bot filtering in a CDN or a hosting firewall is not, and it blocks AI agents silently. Check response codes per user agent before you conclude that your content is the problem.
Treat crawl frequency as a leading indicator. A fetch proves access, an answer proves relevance, and only the panel measures the second one.
Benchmark against the brands AI actually names
The competitor set inside AI answers rarely matches the one in your SERP report. Review platforms, directories, marketplaces and forum threads take a large share of citation slots, because they answer comparison questions in a format a model lifts cleanly.
So record every brand and every domain named per prompt, then count. A ranked competitor table across three cycles tells you two things a ranking report never does: Which companies own your category in the eyes of a model, and which third party pages a model trusts more than your own.
The fix follows the finding. If a directory listing holds the slot, the work is a placement and a consistent profile, not another blog post. If a competitor’s own page holds it, compare the two openings sentence by sentence and rewrite yours to answer first. If a comparison page you do not control holds it, that page is now both a link target and a PR target. Backlinks and third party E-E-A-T signals decide which of two equally clear answers gets picked, and a clean entity definition through structured data decides whether a model is sure which company you are.
A monthly report that holds up to a second look
Cadence beats precision here. Run the panel monthly, judge it quarterly. AI systems get retrained and re-ranked on their own schedule, so a single month that moves five points is usually noise, while three months moving in one direction is a trend.
Six blocks make an AI visibility report defensible:
- The five numbers, with the delta to the previous cycle and to the baseline.
- The competitor table, ranked, with movement marked.
- Prompts won and prompts lost since the last run, named individually.
- Your cited URLs, including the ones that dropped out.
- Generative AI impressions from Search Console, as the one external check on your own panel.
- What was shipped during the period, so results can be attributed to work rather than to luck.
Set the baseline before you change anything. A panel first run after the initial fixes has no starting point, and every later number becomes an assertion. That missing baseline is the most common gap we find when we take over an existing project as part of ongoing SEO support.
Give it 90 days before you judge the result. Pages reached by a live search layer can appear in answers within weeks, while the training side of a model only shifts when it is retrained, on a schedule nobody outside the vendor controls. A visibility index for classic search behaves the same way: One reading is a data point, four are a direction.
Five measurement mistakes that manufacture progress
- One system, one conclusion. ChatGPT, Google AI Mode, Perplexity and Copilot draw on different indexes and different training data. A panel run only where you already do well produces a flattering number and no information.
- A panel of five prompts. Too small to separate a change from noise, and in practice always the five questions where the brand looks best.
- Logged in sessions. Memory and personalization feed your own history back to you. Every run starts clean or the whole cycle is worthless.
- A single vendor score in the board deck. Report the presence rate and the citation rate underneath it, or nobody can check the claim, including you.
- No baseline and no change log. Without both you can show a rising line and still not know what caused it, so you cannot repeat it.
All five make the number easier to produce and impossible to defend. Measurement that cannot survive a critical question is worse than none, because it directs budget.
We set up the panel, run the baseline, report the five numbers every month and work the competitor list until the citation slot is yours. See how we work in our case studies.
Frequently asked questions about AI search visibility
How do I measure AI search visibility?
Run a fixed set of 20 to 40 customer questions across ChatGPT, Google AI Mode and Perplexity every month, and record for each answer whether you are named, whether one of your URLs is cited, and which competitors appear instead. Your presence rate, the share of prompts that name you, is the headline metric.
Can Search Console show AI Overviews data?
Yes, since June 2026. The Search Generative AI performance report shows how often links to your site were shown inside AI Overviews and AI Mode, grouped by page, country, date and device. It reports impressions only, with no clicks, click through rate or position, and the rollout does not yet cover every property.
How do I track AI visibility across several systems?
Use one panel and run it identically everywhere: Same prompts, same order, fresh sessions, three repetitions per system. Store the results per system so you can see where you are strong. Tools automate the running; the design of the panel still decides whether the data means anything.
Why do competitors show up in AI search and I do not?
Usually one of three reasons: A crawler cannot reach or render your page, your page never answers the question in one liftable sentence, or a third party page such as a directory or review site answers it better and holds the citation slot. Check crawler access first, because a blocked page invalidates everything else.
How long does it take to improve AI visibility?
Plan for 90 days, which is three panel cycles. Pages reached by a live search layer can appear in answers within weeks, while the part of a model’s knowledge that comes from training only shifts when that model is retrained.
What counts as a good presence rate?
There is no universal benchmark, because the number depends entirely on how competitive your prompt set is. Use your own first run as the baseline and your competitor table as the ceiling. A brand named in half of its category prompts sits strong; a brand named in none has an access or clarity problem rather than a scoring problem.
How do you calculate share of voice in AI search?
Divide your brand mentions by every brand mention recorded across the panel in the same cycle. If 270 answers hold 690 mentions and 113 of them are yours, your share of voice is 16 percent. Fix the counting rule first: Count each brand once per answer rather than once per occurrence, or one wordy paragraph outweighs four separate answers.
How many prompts does an AI visibility panel need?
Twenty is the floor and thirty is the working number. At 30 prompts and a presence rate near 42 percent, the 95 percent band runs roughly 18 points either way, so only large moves or a consistent direction across three cycles mean anything. Going to 120 prompts halves that band, which pays off only if you also stop reading single months.
Sources
- Google Search Central: “AI features and your website“, Google, 2026.
- Google Search Central Blog: “Introducing Search Generative AI performance reports in Search Console“, Google, June 2026.
- Search Console Help: “Generative AI performance report (Search)“, Google, 2026.
- Search Console Help: “What are impressions, position, and clicks?“, Google, 2026.