AEO Prompt Tracking: How to Choose, Run and Read the Prompts That Measure AI Visibility

An AEO prompt is a question you run through AI answer engines to see whether your brand is named or cited in the answer. How to choose AEO prompts, how many to track, how often to run them, and which of the resulting numbers hold up, with every source linked and read on 10 October 2026.

By TopCitedPublished
Back to Blog

An AEO prompt is a question you run through an AI answer engine, such as ChatGPT, Gemini, Perplexity or Google's AI Overviews, to see whether your brand shows up in the answer: named, linked as a source, or missing. AEO prompt tracking means running a fixed set of those prompts on a schedule and watching how the answers change. It is the closest thing answer engine optimization has to rank tracking. It is also easy to get wrong, because the prompts you choose decide most of what the numbers will say.

This guide covers how to choose AEO prompts, how many to track, how often to run them, and which of the resulting numbers you can trust. It draws on three published sources, each linked where it is used and each read on 10 October 2026:

HubSpot sells an AEO product, and one of the SparkToro researchers works at an AI tracking company. TopCited also sells monitoring. So we say who said what, and we quote rather than paraphrase where the wording matters.

Key definitions

TermWhat it means
AEO promptA question, written the way a real person would ask it, that you run through answer engines to check whether and how your brand appears.
Prompt set (or prompt library)The fixed list of AEO prompts you track. Every rate below is calculated against it, so it is the denominator of everything you report.
Discovery promptA prompt that does not name your brand, such as "best tools for X". It shows whether you are in the answer at all.
Branded promptA prompt that names you, such as "what is [your brand]". It shows whether the answer describes you correctly.
Mention rateThe share of tracked prompts, or of runs, where the answer names your brand, with or without a link. Also called visibility.
Citation rateThe share where the answer links to one of your URLs as a source.
Share of voiceYour brand mentions divided by all brand mentions, yours and competitors', across the same prompt set.

Two things people mean by "AEO prompt"

  • A prompt you track. You run the question through answer engines to measure your visibility. This guide is about this kind.
  • A prompt you write with. You give a chatbot the prompt to draft answer-style content. These can be useful, but they are writing aids. They measure nothing, and the advice below does not apply to them.

Why the prompt set decides the result

Three properties of answer engines make AEO prompt tracking harder than keyword rank tracking.

There is no prompt volume data. PostHog's handbook puts it plainly: "There's no volume data for prompts, so we can't tell whether anyone actually asks the questions we track" (PostHog). Search tools tell you how often a keyword is searched, and nothing equivalent exists for what people type into a chatbot. Every prompt set is therefore someone's assumption about what real people ask.

The same prompt gives different answers. In the SparkToro study, 600 volunteers ran 12 prompts through ChatGPT, Claude and Google's AI Overviews or AI Mode a combined 2,961 times, in November and December 2025. Two findings:

  • The lists almost never repeat. There was a less than 1 in 100 chance that ChatGPT or Google's AI, asked 100 times, would give the same list of brands in any two responses.
  • The order repeats even less. It took "more like 1 in 1,000 runs" before two lists appeared in the same order (SparkToro and Gumshoe.ai).

PostHog sees the same thing in its own tracking: "the same question can produce a different answer depending on the model, the phrasing, and sometimes just the run."

People phrase the same need very differently. SparkToro asked its volunteers to write their own prompts for a single intent: choosing headphones for a family member who travels. The 142 prompts they wrote had an overall semantic similarity of 0.081. Rand Fishkin's summary is that people "don't reduce their search intent to the fewest, most logical set of 2-5 keywords, they get creative, and weird, and highly specific."

PostHog's handbook draws the consequence: "prompt selection is the result. Pick prompts where you're strong and you win. Pick prompts where a competitor is strong and they win."

The study has a reassuring half too. The variation is mostly in the order and the exact list, not in which brands are considered at all:

  • Across differently worded prompts. The 142 headphone prompts produced 994 responses, and brands such as Bose, Sony, Sennheiser and Apple showed up in 55-77% of them.
  • Across repeats of one prompt. In ChatGPT's answers about cancer care hospitals on the US West Coast, City of Hope appeared in 69 of 71 answers, a 97% visibility rate. It was the top mention in only 25 of them.
  • 2,961
    AI answers collected for the SparkToro and Gumshoe.ai consistency study
  • <1 in 100
    Chance that ChatGPT or Google's AI gave the same brand list in any two of 100 runs
  • 0.081
    Overall semantic similarity of 142 human-written prompts for one intent
  • 97%
    City of Hope's visibility across 71 ChatGPT answers, while it was top mention in only 25

That is the whole case for AEO prompt tracking in one study: the answers are unstable, but whether a brand appears across many answers is a stable enough signal to measure.

How to build an AEO prompt set

  1. Start from the topics that matter commercially. List the handful of subjects you most need to be visible for, by product or service line. PostHog organizes its set "by product and topic, so we can see where we're weak rather than just getting one blended figure". It also weights the set toward prompts "where someone is choosing a tool", such as "best session replay tools" or "X vs Y".

  2. Seed prompts from evidence that people ask them. With no prompt volume to go on, borrow signals from elsewhere:

    • Search demand for the matching keywords. PostHog calls Google search volume "the only 'volume' data that exists", while warning that it is imperfect.
    • Your own Search Console queries.
    • Sales calls and support tickets.
    • What new customers say they asked. PostHog's onboarding form asks what people were doing when they found it, "including the prompts they remember using".

    HubSpot's guide seeds its list from personas, customer journeys and pain points.

  3. Write them the way people ask. People phrase things to an LLM "longer, more conversational, with more context about their situation" than they do in a search bar (PostHog). A keyword-style prompt such as "crm small business" is not what anyone types into a chatbot. Write a full question with a situation and a constraint, for example: "What CRM would you recommend for a five-person agency that already uses Google Workspace?"

  4. Write several phrasings per intent. Real phrasing varies so much that one "ideal" prompt per intent is a sample of one. SparkToro's Amanda Natividad frames the real question as "am I showing up reliably across the full semantic neighborhood of this intent?" (SparkToro). Group the phrasings so that you report on the intent, not on one sentence.

  5. Keep branded prompts in a separate bucket. Branded prompts check accuracy: does the answer get your pricing, features and positioning right? They should not count toward visibility, because you appear in them by definition. PostHog tracks prompts like "what is PostHog" for exactly this reason and excludes them from its visibility calculation.

  6. Tag every prompt. HubSpot's guide clusters prompts by topic, intent and region, then tags each with a funnel stage: top, middle or bottom. Tags let you say where you are visible and where you are not, instead of quoting one blended percentage.

  7. Remove rigged and unrealistic prompts. Leave out:

    • Prompts whose answer you already know. PostHog's own example of a prompt rigged in its favour is "the best product analytics tool with a fun mascot and a weird name".
    • Prompts with no evidence that anyone asks them.
  8. Size it to what you can maintain. The two sources pull in opposite directions:

    • HubSpot: "Plan for at least 20 to 100 prompts per brand as a starting point." Its guide describes the practice as running 50 to 200 prompts weekly.
    • PostHog: it went the other way and "pruned our prompt set aggressively", because "a smaller set focused on our ICP tells us more than a large set padded with unreliable and/or unrealistic prompts."

    The two are compatible. The count matters less than whether each prompt earns its place.

How to run your AEO prompts

  • Run every prompt more than once. PostHog: "The same prompt can return different answers on different runs, so any single sample is a snapshot rather than a measurement."
    • The study's bar is high. Fishkin's takeaway is that to know an AI's set of recommendations you need to ask "usually at least 60-100X" and average the results. Across a 100-prompt set, that is 6,000 to 10,000 answers per cycle.
    • The practical substitute is several phrasings per intent (step 4), plus repeated runs on the prompts that matter most. Read the results as percentages across each cluster.
  • Run on a fixed schedule, on the same engines. HubSpot's guide describes running the full library on "a weekly or biweekly cadence". Whatever cadence you choose, keep it and the engine list stable. Otherwise you cannot tell a real change from a change in method.
  • Record three things per answer: whether you are named, whether one of your URLs is linked, and who else appears. That is PostHog's definition of a tracked prompt: one run "on a regular schedule, recording whether we get mentioned, whether we get cited, and who else shows up."
  • Review the set on a cycle. HubSpot recommends refreshing the library "on a quarterly cycle, with lighter monthly reviews layered in." It also says to investigate any prompt that returns zero citations across all engines for three or more consecutive cycles.
  • Expand with care. HubSpot's guide notes that synthetic prompt generation can take a library that "started with 150 hand-written prompts" to "300 or more". It also warns that unchecked expansion floods a library with redundant prompts that dilute reporting.

Which AEO prompt metrics to trust

MetricWhat it countsWorked exampleHow far to trust it
Mention rate (visibility)Share of answers that name you, with or without a linkCity of Hope in 69 of 71 ChatGPT answers, 97% (SparkToro)The most robust of the five. The SparkToro authors call visibility % across dozens to hundreds of prompts, run multiple times, a reasonable metric.
Citation rateShare of answers that link to one of your URLsCited in 40 of 200 tracked prompts is 20% (HubSpot's example)Useful, and separate from mentions: PostHog reports that the two move independently.
Share of voiceYour mentions divided by all brand mentions, times 10020 of 100 mentions is 20% (HubSpot's example)Meaningful only against the same prompt set over time.
Position in the answerWhere in the list your brand appearsCity of Hope was the top mention in only 25 of 71 answers (SparkToro)Do not rely on it. Two answers came back in the same order about 1 in 1,000 runs.
Click-through rateClicks divided by the times you appearedNone availableNot computable. Outside your own set you do not know how many answers you appeared in, so there is no denominator (PostHog).

Two rules follow from the table:

  • Compare a number with itself over time, not with someone else's number. In PostHog's words, "citation rate is only meaningful against your own tracked set."
  • Change the prompt set and you start a new series. Adding or removing prompts moves every rate, so note the date of each change.

For the metrics side in more depth, see AI search visibility metrics and KPIs.

What the three sources recommend, side by side

SourceHow many promptsHow often to runHow often to revise the setWho is speaking
HubSpot, AEO prompt tracking for marketing teams (Justina Thompson, updated 17 August 2026)At least 20 to 100 per brand to start; describes 50 to 200 run weeklyWeekly or biweeklyQuarterly, with lighter monthly reviewsHubSpot sells an AEO product.
PostHog, AEO handbookNo number; a smaller set focused on its ideal customers, pruned aggressivelyOn a regular scheduleRefined as more first-party data comes inPostHog describing its own practice; it tracks with a tool called Gauge.
SparkToro and Gumshoe.ai study (Rand Fishkin, 27 January 2026)12 prompts in the main study; visibility across dozens to hundreds of promptsUsually at least 60-100 runs per prompt to know an AI's recommendation setNot coveredOne researcher works at Gumshoe.ai, an AI tracking startup, and the post says so. Its authors say they are not professional researchers.

Reading someone else's AEO prompt numbers

You will see charts showing a brand winning or losing AI visibility across a set of prompts. PostHog's handbook suggests two questions for any such claim, its own included:

  • "Who picked the prompts, and is there any evidence people actually ask them?"
  • "Is this compared to itself over time, or to someone else's number?"

Fishkin makes the same point from the other side. Because answers vary from run to run, "if you don't like an answer, or your brand doesn't show up where you want it to, just ask a few more times." A number measured on a prompt set you have not seen, run an unknown number of times, tells you very little.

Common AEO prompt mistakes

  • One prompt per topic. It is a sample of one, in a system where phrasing and repeated runs both change the answer.
  • Prompts written like keywords. Nobody asks a chatbot "crm small business". Write the question a person would actually ask.
  • Branded and discovery prompts in one number. Branded prompts inflate visibility, because you appear in them by definition.
  • Reading position as a ranking. Order is the least stable thing in an AI answer.
  • Comparing your rate with a competitor's from a different set. The rates share a name and nothing else.
  • Never pruning. Prompts nobody asks, and prompts that return nothing cycle after cycle, dilute the figure you report.
  • Treating one run as a result. It is one draw from a distribution.

Where TopCited fits

TopCited calls an AEO prompt a query and a prompt set a query set. TopCited's content monitoring:

  • runs your query sets on a schedule you set (cron, with timezone support)
  • records the brand mentions in the answers
  • reports gap analysis: the queries where competitors appear and you do not
  • reports share of voice against the competitors mentioned alongside you

Everything above about choosing prompts applies to choosing queries. The set is the denominator, so it deserves the most care.

When a gap shows up, the CORE optimizer rewrites your existing content and scores it before and after, so the gap leads to a rewritten page rather than a list of issues. To see where you stand first, request a GEO audit through the form on our site.

For the wider picture of tracking your brand in AI answers, see how to track brand mentions in AI search. For choosing a tool, see the AEO tool guide.

FAQ

Frequently asked questions

An AEO prompt is a question, written the way a real person would ask it, that you run through AI answer engines such as ChatGPT, Gemini, Perplexity or Google's AI Overviews to see whether your brand is named or linked in the answer. A fixed list of them, run on a schedule, is a prompt set, and every visibility rate you report is calculated against it.

There is no agreed number. HubSpot's guide says to plan for at least 20 to 100 prompts per brand as a starting point. PostHog's handbook argues for a smaller set focused on its ideal customers, pruned aggressively. A prompt should earn its place: evidence that people ask it, and a topic you need to be visible for.

Often enough to see a trend, on a fixed schedule. HubSpot describes running the full library weekly or biweekly. Because the same prompt returns different answers on different runs, run each prompt more than once. SparkToro's Rand Fishkin suggests usually at least 60-100 runs to know an AI's set of recommendations for one prompt.

Answer engines generate each response afresh. In the SparkToro and Gumshoe.ai study, there was a less than 1 in 100 chance that ChatGPT or Google's AI would return the same brand list in any two of 100 runs, and about 1 in 1,000 for the same order. Which brands appear at all was much steadier than their order.

Not as a primary metric. Order is the least stable part of an AI answer. In the SparkToro study, City of Hope appeared in 69 of 71 ChatGPT answers but was the top mention in only 25. Mention rate across many prompts and runs is the more reliable number.

No. Keep them in a separate bucket. A prompt that names you will mention you by definition, so it inflates visibility. Branded prompts are for checking accuracy: whether the answer gets your pricing, features and positioning right. PostHog tracks them for that and excludes them from its visibility calculation.

From proxies and first-party evidence: search demand for the matching keywords, your Search Console queries, sales calls, support tickets, and what customers say they asked. PostHog's onboarding form asks new users which prompts they remember using. Write each one as a full question, the way a person would ask it.

Yes, under the name queries. You group queries into query sets, and TopCited's content monitoring runs them on a schedule you set (cron, with timezone support), records brand mentions, and reports gap analysis and share of voice against the competitors mentioned alongside you.

The short version

An AEO prompt is a question you run through answer engines to see whether they name or cite you. There is no prompt volume data, and the same prompt returns different answers, so the prompt set decides the result. To build one that measures something:

  • Build it from evidence that people ask the questions.
  • Write several realistic phrasings for each intent.
  • Keep branded prompts separate.
  • Run each prompt more than once, on a fixed schedule.
  • Read mention rate across clusters, over time.

Ignore position in the answer, and be wary of any number measured on a prompt set you have not seen.