Skip to main content


Prompt Tracking: How to Choose the Prompts That Measure Your AI Visibility

Prompt tracking: A small, deliberate selection of prompts is chosen from many possible questions to measure AI visibility
Dominik Breitbach, founder of taismo

Dominik Breitbach Founder and Lead SEO Strategist at taismo

Dominik Breitbach founded taismo GmbH, an SEO and GEO agency from Munich, and works as lead SEO strategist on SEO retainers and visibility in AI answers. He allocates the prompt tracking slots across client projects himself and knows what a prompt costs per month.

⏱ Reading time: 12 min🔄 Last updated: 22 September 2026

Prompt tracking means you define a fixed list of prompts, run it regularly against ChatGPT, Perplexity, Gemini and Google’s AI answers, and count how many answers name your company. That makes the prompt list the measurement itself. It is also the cost driver: Our own account holds 450 prompts in one tool and 200 prompts across three projects in the other, and every prompt costs again in every measurement cycle. Your choice of prompts therefore decides the result before the first model answers.

👉 Want to understand first what AI systems look for? Our generative engine optimization services page explains the groundwork.

What prompt tracking counts

Prompt tracking is the practice of running a fixed set of prompts against AI systems at regular intervals and recording which brands and sources appear in the answers. It is the measurement layer of answer engine optimization (AEO), the discipline that is also called generative engine optimization or GEO. The method has three steps: You define the prompt set, you run it against several models, and you count what the answers contain.

Three metrics come out of every run, and each one answers a different question:

  • Mention rate: The percentage of answers that name your brand. This number measures your presence as an entity and depends on brand mentions across the whole web.
  • Citation rate: How often the model links one of your URLs as a source. This number measures your pages, and you can break it down to the single URL.
  • Share of voice: Which other brands appear in the same answers, and how often. Share of voice turns a bare percentage into a position in your market.

All three metrics share one condition: They describe the prompts you asked and nothing else. A run over 15 prompts is a statement about 15 prompts. That is the core difference to classic rank tracking. On Google you track a query and a results list. In AI search you track a question and a written answer that changes every time you ask again.

Why the prompt list decides the result

Two companies with identical market positions can report completely different AI visibility, simply because they track different prompts. The effect is large enough to make any visibility number worthless without its prompt list.

Take an example from our own field. The prompt “What is answer engine optimization?” gets an answer from every major model that names no agency at all. The prompt “Which agency can help me get mentioned in ChatGPT?” almost always returns three to five names. A company that tracks only definition prompts lands at a mention rate close to zero and concludes that AI visibility is out of reach. A company that tracks only its own brand name lands at 90 percent and believes it leads the market. Both numbers are calculated correctly, and both are useless for decisions.

Google adds a technical reason. Its documentation on AI features explains that AI Overviews and AI Mode may use a “query fan-out” technique, issuing multiple related searches across subtopics and data sources to build one response. The prompt a person types is only the trigger. You never see which sub-queries the system actually resolved, a mechanism we explain in the glossary entry on query fan-out. That makes the one decision you control even more important: Which triggers you track.

FAQs build content, prompts measure it

In client projects we hear the same suggestion again and again: The website already has hundreds of FAQ questions, so why not upload them as the prompt list? The confusion makes sense because both are questions. They do two different jobs.

An FAQ is a question you ask so that your own page answers it. It belongs in the visible copy and in your schema markup, it raises your chance of being cited, and it costs nothing to run. A prompt is a question a person asks a model. It belongs in the tracking plan, it occupies a paid slot, and it answers nothing for your visitors.

The scale makes the difference concrete. Our own site covers 2,445 indexed URLs, and depending on how you count, they hold between 2,400 and 6,100 FAQ questions (as of 11 September 2026). The largest plan we found on the market covers 1,000 prompts. Even that plan could track only a fraction of our own FAQ inventory, and most of those questions never trigger a brand name in any model.

Which prompts to track: The two-condition test

We give a prompt a tracking slot only if two conditions are both true. First, the answer carries a vendor or buying decision. Second, the answer distinguishes between vendors. If either condition fails, the model names no brand in its answer and the paid slot burns.

Four prompt types pass the test reliably:

  • Selection prompts (“Which tool is best for X?”) force the model to produce a list of names.
  • Comparison prompts (“X or Y, which fits a company with 80 employees?”) add the reasoning the model uses to position you.
  • Brand prompts (“What does company X do?”, “Is X a good choice for B2B?”) show whether the model knows you at all and which competitors it names in your place.
  • Location prompts (“Who offers X in Munich?”) belong in every local plan, because AI answers name especially few providers per city.

Three prompt types fail the test even though they look relevant: Pure definition prompts without vendor context, service questions about your own operations (“How long does delivery take with you?”) and prompts with a trivial factual answer. All three belong in the visible copy of your website, where they build citability. In a tracking plan they only cost money.

When a prompt earns a slot in prompt trackingWhen a prompt earns a tracking slotBoth conditions must be true1. The answer carries a vendor or buying decision.2. The answer distinguishes between vendors.Tracking slotthe model names brands hereSelection promptsWhich X is best for a 200-person company?Comparison promptsX or Y for a mid-sized manufacturer?Brand promptsWhat does company X do? Is X any good?Location promptsWho offers X in Munich?Page copyanswers here name no brandPure definitionsWhat is answer engine optimization?Service questionsHow long does your delivery take?Trivial factsHow many states does the US have?These build citability on your pagesFig. 1 · taismo
Fig. 1: The two-condition test for a tracking slot. Both conditions must be true, otherwise the model names no brand in its answer.
💡
Pro tip: Start with two brand prompts before you spend budget on selection prompts. “What does <your brand> do?” and “Is <your brand> a good choice?” cost two slots and answer the most expensive question first: Does the model know you, and which three competitors does it name when it does not?

What prompt tracking costs: Two quotas

Vendor pricing pages show a prompt count, and that number sounds generous. In daily use you run into two separate quotas with two different bottlenecks. Our numbers come from our own SE Ranking account, measured on 11 September 2026: The AI search add-on unlocks two tools there, the AI Results Tracker and SE Visible, and each one bills against its own quota.

Tool type Quota What it measures Bottleneck
Own visibility tracker 450 prompts your presence for the keywords you already track, including the overlap with your organic rankings prompt count
Market monitor 200 prompts, 3 projects share of voice, competitor brands, sentiment, cited sources per domain and per URL project count

Three rules decide what a tracking plan really consumes:

  • A prompt counts once, no matter how many models it runs against. We verified this on our own project with five models: 20 prompts consumed 20 of 200 slots. More models cost nothing extra and give you a steadier picture.
  • A prompt costs again in every measurement cycle. The number in your plan is running capacity. If you measure monthly, each prompt stays booked for the whole contract.
  • The project count is the harder limit. Three projects cover three brands at the same time. The 200 prompts would only become the problem with far more brands.

The consequence: Tracking slots get allocated. At taismo, clients with an explicit AEO mandate get 25 to 40 prompts, clients with an ongoing SEO retainer get 12 to 15, and everyone else starts without any. How this effort shows up in pricing is broken down in our article on what generative engine optimization costs. If you want the full service behind the tracking, our generative engine optimization services combine the plan, the content work and the monthly reading of the data.

Two quotas, two bottlenecks in prompt trackingTwo quotas, two bottlenecksOur own account, measured on 11 September 2026Own visibility tracker450promptsmeasures you against tracked keywordsLimit: prompt countMarket monitor200prompts3projectsshare of voice, competitors, cited sourcesLimit: project countOne prompt against five models used 1 slot: 20 prompts, 5 models, 20 of 200 usedFig. 2 · taismo
Fig. 2: Two quotas with two different bottlenecks. Adding models costs nothing, the number of projects sets the hardest limit.

Why repetition makes prompt tracking reliable

Language models answer probabilistically. The same prompt, asked twice, returns two different brand lists. This holds across all major systems and on every repeat.

Rand Fishkin (SparkToro) and Patrick O’Donnell (Gumshoe.ai) quantified the effect in January 2026. 600 volunteers ran 12 prompts through ChatGPT, Claude and Google’s AI answers, 2,961 runs in total. The result: The chance that any two answers out of 100 contain the same list of brands is below 1 in 100. For the same list in the same order it drops to roughly 1 in 1,000. Fishkin’s recommendation is to ask 60 to 100 times and average the results, and he calls tracking a “ranking position” in AI tools foolhardy.

For your plan, this has an uncomfortable consequence. Repetition multiplies with the quota: 15 prompts, measured monthly with enough repetitions, add up to a large number of model calls. If you spread the same budget over 150 average prompts measured once, you get a snapshot without statistical weight. How AI systems decide which brands they recommend stays the same either way. Measuring that decision just gets more expensive the wider you spread.

In practice: A short list, repeated often, beats a long list that runs once a quarter. The visibility index of classic SEO works on a similar principle, it simply rests on far more stable raw data.

How many prompts you need for prompt tracking

A mid-sized B2B company needs 12 to 20 well-chosen prompts. That range is enough to cover the decisions your buyers make, and small enough to repeat each prompt often enough for stable percentages. Our own allocation follows that logic: 12 to 15 prompts for SEO clients, 25 to 40 for clients whose mandate centers on AI visibility.

Three factors move the number up or down:

  1. Number of offers: Each service line with its own buyers needs its own selection and comparison prompts. A company with three distinct offers needs roughly three times the selection prompts of a single-offer company.
  2. Number of markets: Each language and each country is a separate answer space. A German prompt and its English equivalent return different brand lists, so each market needs its own set.
  3. Local footprint: Location prompts scale with the cities you serve. Three locations mean three location prompts per service.

The upper limit comes from your quota divided by your cycle count, and from the repetitions your tool runs per prompt. Calculate a full month before you sign, never a single run.

We build your prompt list before we measure

In our GEO audit we first identify the decision questions your buyers really ask and derive the tracking plan from them. You get the list, the reasoning for each prompt and your starting point as a number.

See the GEO audit

What Search Console adds to prompt tracking

Google introduced a dedicated report for generative AI features in Google Search Console on 3 June 2026, and since 31 August 2026 it is available to all websites worldwide. It is the only AI visibility data that comes directly from Google, and it is free.

The report delivers one metric: Impressions. It shows how often URLs from your site appeared in generative AI features on Google Search, which currently means AI Overviews and AI Mode. You can group that metric by pages, countries, devices and dates. Clicks, click-through rate and queries are outside its scope.

So you see that you appear, and you see which pages make it happen. The question that matters stays open: Which question did you appear for? Your own prompt tracking fills exactly that gap. Together the two sources give you the full picture, because Search Console shows the trend across your whole site and your tracking plan explains individual topics.

Google adds a reassurance in its documentation on AI features: There are “no additional technical requirements” for appearing in AI Overviews or AI Mode beyond being indexed and eligible for a snippet. Prompt tracking therefore reveals how well your existing work carries into AI answers. How much traffic AI Overviews actually pass on is covered in our guide to AI Overviews and SEO.

Your own visibility vs. share of voice

The phrase “measure AI visibility” hides two different questions, and each one needs its own type of tool. Buy the wrong type and you get a clean number for the wrong question.

Question 1: Am I mentioned for my topics? This measurement runs against the keywords you already track and adds the overlap with your organic rankings. It answers whether your SEO work carries into AI answers, and it is the right choice for tracking your own development over time.

Question 2: Who gets mentioned in my place, and which sources do the models cite? This measurement looks at the market. It delivers your share of all mentions, the competitor brands in the same answers, the tone of what the models say about you and the cited sources at domain and URL level.

The cited sources are the most valuable output in practice. They show two things at once: Which of your own pages do the work, and on which third-party domains you would need to appear to get into the answers. For the technical side, meaning whether models may read your pages at all, the controls remain your AI crawler rules and your llms.txt. Our broader guide to measuring AI search visibility covers the remaining data sources, from referral traffic to server logs.

A 15 prompt tracking plan in five steps

Fifteen prompts is the size we start with for a client on an SEO retainer. The plan comes together in five steps.

Step 1: Collect decision questions. The best sources sit inside your company. Sales calls, proposal emails and the questions that come up in every first meeting give you real wording. You can spot invented prompts because nobody has ever phrased a question that way.

Step 2: Filter with the two-condition test. Every collected question runs through the two conditions above. In our experience about one third survives.

Step 3: Distribute the 15 slots. Our default split is 5 selection prompts, 4 comparison prompts, 3 brand prompts and 3 location prompts. For a nationwide vendor, the three location slots move to comparison prompts.

Step 4: Set models and cycle. Four to five models cost nothing extra, so you take every model your tool offers. A monthly cycle has proven itself, because model updates and index refreshes become visible at that pace.

Step 5: Freeze the list. The list stays unchanged for at least two cycles. If you swap prompts in between, you end up measuring changes to your own list. From the third cycle on, you replace individual prompts one at a time and document every change.

A 15 prompt tracking plan with a monthly cycleA 15 prompt tracking plan5 selection4 comparison3 brand3 locationNationwide vendors move the 3 location slots into comparison promptsCycle 1baselineCycle 2first changeCycle 3first trendList frozenMonthly, 4 to 5 modelsswap one prompt at a timefrom cycle 3, documentedFig. 3 · taismo
Fig. 3: The default split of 15 tracking slots across four prompt types, plus the monthly cycle with a frozen list.

Four mistakes that break prompt tracking data

We see these four mistakes regularly, and each one on its own makes the numbers unusable.

1. The FAQ list becomes the prompt list

This is the most common mistake and the most expensive one. It burns the quota on questions for which no model ever names a brand. The distinction is explained in the section on FAQs and prompts above.

2. A rank position gets reported as a KPI

“Position 3 in ChatGPT” sounds familiar and has no basis in the data. A written answer has no results list, and according to the SparkToro numbers the order of named brands repeats roughly once in a thousand runs. Mention rate, citation rate and share of voice are the metrics that hold up.

3. One run gets read as a trend

A value from a single run says nothing about development. The first cycle delivers a baseline, the second the first change, and from the third cycle on you can talk about direction. Classic SEO asks for the same patience.

4. Measuring replaces the work

A tracking plan improves nothing by itself. It shows whether the foundations carry: Consistent facts about your company, a clean entity in your structured data and content that answers questions directly. Our free AI visibility checker shows in two minutes where a single page stands on those foundations.

Let’s define your 15 prompts

We tell you which decision questions in your market name brands at all, how many tracking slots you need for them and what already works in your favor today. The first call is free.

Request a free consultation

Frequently asked questions about prompt tracking

What is prompt tracking?

Prompt tracking is the regular measurement of a fixed set of prompts across AI systems such as ChatGPT, Gemini, Perplexity and Google AI Mode. It records which brands each answer names and which sources it cites, and turns that into mention rate, citation rate and share of voice.

How many prompts should you track for AI visibility?

A mid-sized company needs 12 to 20 well-chosen prompts, provided each one runs often enough. The number of repetitions matters more than the number of prompts, because language models answer differently on every run.

Which prompts should you track?

Track prompts whose answer carries a buying decision and distinguishes between vendors: Selection, comparison, brand and location prompts. Definition questions and service questions about your own business belong in your page copy instead.

Can you track AI visibility for free?

Partly. The generative AI performance report in Google Search Console is free and shows impressions from AI Overviews and AI Mode. Mentions in ChatGPT, Perplexity or Claude and comparisons with competitors require a paid tool or manual sampling.

How often should you run prompt tracking?

Monthly is the useful cycle. Measuring more often raises costs without improving the result, and measuring less than quarterly makes you miss model updates.

Is there a rank tracker for AI answers?

AI answers have no stable positions. An answer is running text without a results list, and the order of named brands changes from run to run. Reliable metrics are mention rate, citation rate and share of voice.

Sources

0%
Get insider knowledge first!
taismo Logo

© taismo GmbH

Address


Weißenfelder Str. 6
85551 Kirchheim near Munich, Germany