·
Skip to main content

The GEO journal9 min read

How to Build a Client Report for AI Visibility

A client-facing AI visibility report contains an executive summary, per-engine metrics for ChatGPT, Perplexity, Gemini and Claude, a competitor benchmark, and recommended actions. Send it monthly, built on continuous underlying monitoring. One rule makes the numbers defensible: a locked prompt set, measured the same way every period, so each month compares to the last.

What clients expect the report to answer

A client pays for three answers: are we visible in AI answers, compared to whom, and what changed since last period. Every section of the report should serve one of those three questions. Anything that does not — raw transcripts, tool screenshots, unexplained composite scores — belongs in an appendix.

Clients are asking now because the traffic data has moved. Pew Research Center reported on 22 July 2025 that Google users who encountered an AI summary clicked a traditional search result link in 8% of all visits, while users who did not encounter one clicked nearly twice as often, in 15% of visits. And Gartner predicted in a February 2024 press release that traditional search engine volume will drop 25% by 2026, due to consolidation of consumer traffic across AI chatbots and other virtual agents. A client who has seen either number wants to know where their brand stands in the channel absorbing that attention.

Open with an executive summary: three or four sentences, in plain non-technical language, stating the headline visibility change and the main reason behind it. Write it so the client can forward it internally without edits, because that paragraph is the part their management will actually read.

Run two cadences and explain the difference to the client. Monitoring is continuous, because AI answers shift from day to day and a single check is not a measurement. The client-facing report is monthly, because a month of accumulated runs turns daily noise into a trend the client can act on. Continuous data underneath, monthly narrative on top.

Finally, define in month one what showing up in ChatGPT means for this client: the brand named in the answer text (a mention), a link to the brand's domain among the sources (a citation), or both. The two move independently — see the difference between an AI mention and an AI citation — and a report that blends them cannot explain its own numbers.

The metrics that belong in the report

Report each engine separately. Profound published an answer-engine citation study in July 2025 finding that nearly 89% of AI citations come from completely different sources depending on which model users query. A blended score averages away exactly the differences the client needs to see; a per-engine scorecard shows where the brand is strong and where the work is.

Mentions and share of answer

Count how many tracked prompts mention the brand, per engine, and note where in the answer the mention sits — a first-sentence recommendation is not the same as a name in a closing list. Then compute share of answer, the brand's presence relative to every brand named across the tracked prompts; the method is in how to calculate share of answer. Record mention-without-citation cases separately: they show the engine knows the brand but is relying on third-party descriptions of it.

Citations and sources

For each engine, list the domains cited when the brand comes up, split between owned sources (the client's site) and third-party sources (press, directories, review platforms). Perplexity is the most explicit engine here: its developer documentation, verified on 8 September 2026, describes how the model inserts numbered references in the answer text, with the corresponding source URLs delivered alongside. That explicitness makes Perplexity citations the easiest per-engine KPI to defend in front of a client, since every count traces back to a visible source list.

Sentiment and accuracy

Whether the brand is mentioned is half the metric; how it is described is the other half. Add a sentiment line per engine — positive, neutral or negative framing — and an accuracy check on whether the engine describes the offer, pricing model and locations correctly. Consider a hypothetical answer that recommends a brand while attaching an outdated description of it: the mention would help less than it appears to.

Use this checklist to record each engine's answers. Check citation availability in each run; the table is a reporting worksheet, not a summary of Namedrop test results.

EngineMentionsCitations
ChatGPTRecord whether and where the answer names the brandSave any cited URLs; record their absence when none are provided
PerplexityRecord whether and where the answer names the brandMatch numbered references to source URLs, as described in Perplexity's documentation
GeminiRecord whether and where the answer names the brandSave any cited URLs; record their absence when none are provided
ClaudeRecord whether and where the answer names the brandRecord whether web search was used and which URLs were cited

The prompt set and the competitor benchmark

The trend line is only as good as the prompt set's stability. Agree a set of prompts with the client — the questions their buyers actually ask — and lock it. Changing prompts each cycle destroys the trend: month two no longer measures the same thing as month one, and every movement becomes unexplainable. When the business changes, add prompts with a marked start date and report them separately until they have a history.

Run each prompt several times per period, never once. OpenAI's documentation on the seed parameter, verified on 8 September 2026, states that determinism is not guaranteed and points to the system_fingerprint response parameter for monitoring backend changes. If the model's own vendor does not promise identical outputs for identical inputs, a single run is an anecdote; repeated runs across the month are a measurement.

Benchmark against a fixed competitor set on the same locked prompts, never on ad hoc queries. The client question is compared to whom, and on which prompts — so present the benchmark as a per-engine view of the brand against each named competitor on the shared prompt set, not as one abstract composite score. A score without named competitors and named prompts cannot be challenged, which also means it cannot be trusted.

For multi-location and local-service clients, build the set in two layers: a shared core of brand and category prompts, plus a per-location template — the same buyer question phrased for each city — instantiated per market. Aggregate results by location in the report, so a local manager sees their own line while the executive summary keeps the overall picture.

The report template, section by section

The skeleton below is the whole report, in reading order. The executive summary comes first, and every section after it should be answerable in one glance: the number, last period's number, why it moved. The example lines use Alpencab, an invented airport-transfer brand with realistic but fictional values.

Snapshot header. Client, period, engines covered, prompt-set version, run schedule. Example: Alpencab, August 2026, ChatGPT, Perplexity, Gemini and Claude, prompt set unchanged since March, every prompt run weekly.

Executive summary. Three or four plain sentences. Example: Alpencab is now mentioned in roughly half of tracked prompts on Perplexity, up from about a third last month; the gain follows the airport-guide pages published in July, which Perplexity now cites. ChatGPT still describes an outdated pricing model, so correcting the source pages is next month's priority.

Per-engine scorecard. One block per engine: mention rate, share of answer, owned versus third-party citations, sentiment — each beside last period's value. Semrush's reporting guide, consulted on 8 September 2026, arranges metrics in tiers based on how close they are to business value; apply the same logic and lead with what the client's management cares about.

Competitor benchmark. The fixed competitor set on the same prompts, per engine. Example: on category prompts, Gemini recommends TransferPro ahead of Alpencab in most runs, citing two review platforms where Alpencab is absent.

Cited sources. The domains each engine relied on this period, owned versus third-party, with new entrants flagged. Example: a regional travel blog entered Perplexity's citations this month, and its description of Alpencab is accurate.

Recommended actions. For every metric that moved, one what-changed-and-why line linking the movement to actions taken during the period, then the next actions with owners. This section answers the client's real question: which of the things they paid for worked.

When two tools disagree, and what to keep as evidence

Two tools tracking the same brand will report different numbers, and neither has to be wrong. They differ in prompt phrasings, run counts, run dates, model versions and, most often, in what counts as a mention — some count any appearance of the brand name, others only a recommendation. Reconcile in that order: compare definitions first, then the prompt sets, then run counts and dates. Most discrepancies dissolve at the definition step.

This is where the continuous monitoring under the monthly report earns its place. When the client questions a number, pull the individual runs behind it and show which answers, on which dates, produced the figure. A report that cannot be decomposed into raw answers is an assertion, not a measurement.

Keep the evidence deliberately. For every figure in the report, store the raw answer transcript, the date and time of the run, the engine, and the model version where the API exposes it. Then set a retention period and delete on schedule: transcripts can contain personal data, and CNIL's guidance on retention, verified on 8 September 2026, states that personal data cannot be kept indefinitely — the data controller must define a retention period tied to the purpose the data was collected for. A workable rule: keep transcripts while the figures they support sit in a live report, plus the contractual dispute window, then purge.

Where to start this week

Sequence matters more than speed here, because the one thing that cannot be fixed later is a prompt set that changed mid-engagement.

First, agree the prompt set with the client before any measurement: a couple of dozen prompts drawn from how their buyers actually ask, plus the per-location template if they operate in several markets. Get written sign-off — that sign-off protects the trend line later.

Second, run the multi-engine baseline with repeated runs spread across the week, and store everything: transcripts, dates, engines, model versions.

Third, fill the template once and label it the baseline. Month one establishes the numbers, month two establishes the trend — resist reading movement into a single period.

Fourth, schedule the monthly cadence and set up the underlying monitoring so runs accumulate between reports; how to automate AI visibility checks covers the setup.

Run a free Namedrop scan to generate the baseline numbers for your first client report.

Sources

Frequently asked questions

How often should I send an AI visibility report to a client?
Send a client-facing report monthly, built on monitoring that runs continuously underneath. AI answers vary from day to day, so daily numbers are noise; a month of accumulated runs produces a trend a client can act on. The continuous layer matters when a number is challenged: it lets you pull the individual answers behind any figure. Weekly reporting is rarely worth the overhead unless a launch or a crisis is in progress.
How do I track brand mentions in Perplexity for a client KPI?
Run the locked prompt set in Perplexity on a fixed schedule and count answers that name the brand, separating mentions from citations. Perplexity's documentation describes numbered references in the answer text tied to source URLs, so every citation count can be traced to a visible source list. That traceability makes Perplexity the easiest engine to defend as a KPI: when the client asks where a number came from, you show the answers and their sources.
What metrics should an AI visibility report include beyond citation counts?
Add share of answer — the brand's presence relative to all brands named on the tracked prompts — plus position within the answer, mention-without-citation cases, the split between owned and third-party cited domains, and sentiment and accuracy of how each engine describes the brand. An engine can recommend a brand while describing it incorrectly, and it can mention a brand without citing it; each situation calls for a different action, so each needs its own line in the report.
Why do two AI visibility tools show different numbers for the same brand?
They measure differently: prompt phrasings, run counts, run dates, model versions and, most often, the definition of a mention vary between tools. Model outputs are also not deterministic, so even identical setups diverge across runs. Reconcile in order: compare each tool's definition of a mention first, then the prompt sets, then run counts and dates. Most gaps close at the definition step; the rest usually trace to sampling differences rather than to an error.
How do I compare my AI visibility against competitors in the report?
Fix a competitor set with the client and measure it on the same locked prompt set, per engine — never on ad hoc queries. Present a table of your brand against each named competitor, showing mention rate and share of answer, and note which sources the engines cite for competitors that you are absent from. That answers the question of compared to whom, and on which prompts, where a single composite score cannot be verified or acted on.