# How to Compare Your AI Visibility Against Competitors

Author: Quentin Megevand
Published: 2026-09-17
Source: https://getnamedrop.ai/en/blog/how-to-compare-your-ai-visibility-against-competitors

A valid competitor comparison in AI search needs three things: the same locked prompt set run on the same engines for every brand, enough repeated runs to separate a real gap from normal answer variance, and raw metrics — mention rate, citation rate, share of voice — rather than proprietary scores that only make sense inside one tool.

## What a competitor comparison in AI search actually measures

A competitor comparison in AI search tracks three things across a fixed prompt set: how often each brand is mentioned in the generated answers, how often each brand's pages are cited as sources, and the resulting share of voice — each brand's slice of all brand appearances across the set. Mention tracking tells you whether the engine considers a brand at all. Citation tracking tells you which pages earn a place in answers, and it is what makes the comparison actionable: a citation points to a specific URL you can strengthen, and a competitor's citation points to a page you can study.

Cover the surfaces where buyers actually ask: ChatGPT, Perplexity, Gemini, and Google's two answer surfaces, AI Overviews and AI Mode. Track them separately, because they select sources differently and a brand can dominate one while staying absent from another.

Google rankings are not a proxy for any of this. In a study published on 11 August 2025, Ahrefs reported that in a dataset of 15,000 prompts, on average only 12% of links cited by ChatGPT, Gemini, and Copilot appear in Google's top 10 results for the same prompt. A rank-tracking report therefore says little about who an assistant names or cites; you have to measure the answers themselves. The stakes are direct: when a buyer asks an assistant to shortlist vendors, a brand absent from the answer leaves the consideration set before any click happens.

## Choose the competitor set and lock the prompt set

The comparison has two inputs. Build both deliberately, then stop touching them for the measurement period.

For the competitor list, start with the brands you would name yourself, then run your category prompts once and note which brands the engines actually surface. The two lists rarely match, and the second one is the market as the engines present it, so include both. Plan for categorical competitors too: engines often answer a niche prompt with a generalist — a broad suite recommended where you expected specialist tools — and that displacement is itself a finding worth tracking. When the engine's framing of your category differs from yours, your comparison must reflect the engine's framing, because that is what buyers see.

For the prompt set, work from the buying questions of the category: what a prospect asks before choosing, from broad comparisons to specific use cases. Mix brand-neutral prompts, which name no vendor, with brand-adjacent prompts such as alternatives to a named competitor. Fifteen to twenty prompts is enough to start.

Then lock the set and date it. Adding or removing prompts mid-period invalidates the trend, because next month's share of voice would be computed over a different denominator. If the category shifts, version the set instead: retire the old list, date the new one, and restart the baseline from there.

## How many runs before a difference is real

### Run-to-run variance is the baseline, not a bug

LLM answers are non-deterministic. The same prompt, on the same engine, on the same day, can produce answers that name different brands and cite different sources. A single run per prompt therefore measures noise, not position: one answer naming brand A but not brand B is compatible with both brands appearing about half the time over repeated runs.

Variance also differs by engine. Qwairy's citation study, published on 15 October 2025, analyzed 118,101 AI-generated answers with 669,065 citations across 8 major providers during Q3 2025, comparing how providers such as ChatGPT and Perplexity cite sources. The measurement consequence: how many sources an engine cites per answer sets the denominator of your citation rate, and an engine that lists many sources gives long-tail pages more chances to appear, which makes run-to-run results noisier. Measure each engine against its own baseline and never pool answers across engines.

Our working rule: run each prompt five times per engine per measurement period, spaced across different days rather than back to back. Five runs is the point where always mentioned, sometimes mentioned and never mentioned become distinguishable states; one run cannot tell the second from the third. Score mention rate as the fraction of runs naming the brand, treat gaps smaller than one run in five as ties, and let the day-to-day spacing absorb model and index updates instead of freezing one moment.

### Separating a caused improvement from normal fluctuation

When your mention rate rises after a content change, three explanations compete: the change worked, the engine moved for everyone, or you are reading noise. Hold the prompt set constant, so the denominator cannot explain the shift. Compare against the competitor baseline over the same window: if every brand rose, the engine changed, not your position. And require persistence: a real improvement holds across at least two consecutive measurement periods, while fluctuation reverts. A one-period spike, however welcome, is not evidence.

## Metrics that compare across tools, and scores that don't

Four raw metrics survive contact with any tool, because you can recompute them by hand from archived answers:

- Mention rate: the share of runs in which the brand is named in the answer.
- Citation rate: the share of runs in which one of the brand's URLs appears among the cited sources.
- Share of voice: the brand's mentions as a share of all tracked brands' mentions over the prompt set; we explain the manual calculation in [how to calculate share of voice in AI search by hand](/en/blog/what-is-share-of-voice-in-ai-search-how-to-calculate-it).
- Average position: where in the answer the brand appears, because a first-sentence recommendation carries more weight than a closing mention.

Vendor scores are a different object. Ahrefs' help documentation, consulted on 17 September 2026, defines the AI share of voice in Brand Radar as a brand's percentage share of impressions compared to other tracked brands — a clear definition, but computed over Ahrefs' own tracked-brand universe. Semrush's knowledge base, consulted the same day, describes its AI Visibility Toolkit data as covering ChatGPT, Gemini, Google AI Overviews, and AI Mode. SE Ranking's product page, also consulted on 17 September 2026, states that its AI Search Toolkit analyzes answers generated by ChatGPT, Perplexity, Gemini, AI Mode, and AI Overviews. Different engine sets, different prompt universes, different aggregation: a Rankscale AI readiness score, a Semrush visibility score and an Ahrefs share-of-voice figure cannot be placed on one axis, because no shared methodology connects them. None of these tools is wrong; they answer differently posed questions. Engine coverage and pricing tiers also change frequently, so published third-party tool comparisons go stale within months — verify against the vendor's own documentation before deciding.

| Property | Raw metrics | Proprietary composite scores |
|---|---|---|
| Examples | Mention rate, citation rate, share of voice, average position | Ahrefs Brand Radar AI share of voice, Semrush AI visibility score, Rankscale AI readiness score |
| Recomputable by hand | Yes, from archived answers | No, methodology unpublished or partial |
| Comparable across tools | Yes, same definition everywhere | No, each vendor's own formula and engine set |
| Best use | Cross-tool benchmarking and client reporting | Trend tracking inside a single tool |

The rule for multi-tool situations: pick one tool as the system of record for trends, and when a client brings numbers from another tool, move the conversation to raw counts both sides can verify against saved answers.

## Manual checks, DIY automation, or a monitoring tool

The manual method costs nothing but time: run each prompt in a logged-out browser or a clean profile so personalization does not skew answers, screenshot each answer, and record mentions and citations per brand in a spreadsheet. It is auditable and it works. It stops scaling fast, though: a modest prompt set, rerun five times on two engines, already produces a few hundred answers to read every period.

DIY automation is where platform terms come in. Consumer AI products commonly restrict automated or programmatic access to their interfaces and the extraction of output at scale, and scripted querying of a consumer interface also invites rate limiting and account suspension. Read the current terms of use of each platform before scripting anything against ChatGPT, Perplexity or Google's AI surfaces; the terms change, and the risk is yours, not your tool vendor's. Where the vendor offers an official API, API-based collection is the compliant DIY route: you pay per call, stay inside the developer terms, and get responses that are easier to parse. Keep in mind that API answers and consumer-interface answers are not guaranteed to be identical, so choose one collection method and hold it constant. We describe the practical setup in [how to automate AI visibility checks](/en/blog/can-you-automate-ai-visibility-checks-here-s-how).

Monitoring tools do the collection for you, and several offer free tiers or trials, so cost is not the barrier to starting. The questions that matter are whether the tool covers the engines your buyers use and whether it exposes the raw counts described above, not just a composite score.

## Where to start this week

The first measurement takes a few hours and should already follow the method, so it counts as a baseline rather than a throwaway. In order:

- List three to five competitors: the brands you name yourself, plus any brand or generalist the engines surfaced when you tested your category prompts.
- Write fifteen to twenty prompts from the category's buying questions, mixing brand-neutral and brand-adjacent phrasings, then lock the list and date it.
- Run each prompt five times on at least two engines, spaced across the week, in a clean browser profile.
- Record, per run, which brands are mentioned, which URLs are cited, and where in the answer each brand appears.
- Compute mention rate, citation rate and share of voice per brand and per engine.
- Date and archive the snapshot: the answers, the spreadsheet and the totals.

The output of week one is a baseline, not a verdict. The comparison becomes meaningful at the second dated snapshot, when the locked prompt set and the rerun discipline let you attribute movement instead of guessing at it. From that point the same table becomes the core of [building a client report for AI visibility](/en/blog/how-to-build-a-client-report-for-ai-visibility).

[Run a free Namedrop scan](/en?src=blog-article#scanner) to see where your brand currently stands next to competitors in ChatGPT, Perplexity, Gemini and Claude, and use it as the first dated snapshot of your comparison.

## Sources

- [Ahrefs, Only 12% of AI Cited URLs Rank in Google's Top 10 for the Original Prompt, consulted 2026-09-17](https://ahrefs.com/blog/ai-search-overlap/)
- [Qwairy, Perplexity vs ChatGPT: AI Citation Study (Q3 2025), consulted 2026-09-17](https://www.qwairy.co/blog/provider-citation-behavior-q3-2025)
- [Ahrefs Help Center, AI Visibility Metrics, consulted 2026-09-17](https://help.ahrefs.com/en/articles/15501968-ai-visibility-metrics)
- [Semrush, Where does the data in Semrush's AI Visibility Toolkit come from?, consulted 2026-09-17](https://www.semrush.com/kb/1607-semrush-ai-visibility-data)
- [SE Ranking, AI Search Visibility Tool: Optimize for AI Search, consulted 2026-09-17](https://seranking.com/ai-visibility-tracker.html)

## Frequently asked questions

### What is a good AI visibility score?

There is no universal benchmark, because each tool computes its score with its own methodology, engine coverage and prompt universe, so a score from one tool cannot be compared with a score from another. The meaningful benchmark is relative: your share of voice against named competitors on a fixed, dated prompt set, measured the same way each period. If that share grows across consecutive snapshots while the prompt set stays locked, your visibility is improving, whatever the composite score says.

### What tools track brand mentions across ChatGPT, Gemini, and Perplexity?

Dedicated AI visibility monitors such as Namedrop track whether a brand is cited in ChatGPT, Perplexity, Gemini and Claude and why competitors are cited instead. Established SEO suites have added AI tracking: Semrush's AI Visibility Toolkit draws on ChatGPT, Gemini, Google AI Overviews and AI Mode data, and SE Ranking's AI Visibility Tracker analyzes ChatGPT, Perplexity, Gemini, AI Mode and AI Overviews. Whichever you choose, insist on raw mention and citation counts you can verify by hand, since composite scores are not comparable across tools.

### How many times do I need to run a prompt to get a reliable comparison?

Run each prompt at least five times per engine per measurement period, spaced across different days. Answers are non-deterministic, so a single run measures noise: it cannot distinguish a brand that is always mentioned from one that appears half the time. Five runs make always, sometimes and never distinguishable outcomes, and spacing the runs captures model and index updates. Treat gaps smaller than one run in five as ties, and confirm any trend across two consecutive measurement periods.

### What is Google's equivalent of ChatGPT?

Gemini is Google's conversational assistant and the closest equivalent to ChatGPT. Google also generates answers inside search itself, through AI Overviews and the fuller AI Mode. A comparison should track the three separately: they are distinct surfaces with different source-selection behavior, and a brand can be cited in Gemini while staying absent from AI Overviews on the same prompt. Treat each as its own engine, with its own baseline and its own run count.
