A client-facing AI visibility report contains an executive summary, per-engine metrics for ChatGPT, Perplexity, Gemini and Claude, a competitor benchmark, and recommended actions. Send it monthly, built on continuous underlying monitoring. One rule makes the numbers defensible: a locked prompt set, measured the same way every period, so each month compares to the last.
What clients expect the report to answer
A client pays for three answers: are we visible in AI answers, compared to whom, and what changed since last period. Every section of the report should serve one of those three questions. Anything that does not — raw transcripts, tool screenshots, unexplained composite scores — belongs in an appendix.
Clients are asking now because the traffic data has moved. Pew Research Center reported on 22 July 2025 that Google users who encountered an AI summary clicked a traditional search result link in 8% of all visits, while users who did not encounter one clicked nearly twice as often, in 15% of visits. And Gartner predicted in a February 2024 press release that traditional search engine volume will drop 25% by 2026, due to consolidation of consumer traffic across AI chatbots and other virtual agents. A client who has seen either number wants to know where their brand stands in the channel absorbing that attention.
Open with an executive summary: three or four sentences, in plain non-technical language, stating the headline visibility change and the main reason behind it. Write it so the client can forward it internally without edits, because that paragraph is the part their management will actually read.
Run two cadences and explain the difference to the client. Monitoring is continuous, because AI answers shift from day to day and a single check is not a measurement. The client-facing report is monthly, because a month of accumulated runs turns daily noise into a trend the client can act on. Continuous data underneath, monthly narrative on top.
Finally, define in month one what showing up in ChatGPT means for this client: the brand named in the answer text (a mention), a link to the brand's domain among the sources (a citation), or both. The two move independently — see the difference between an AI mention and an AI citation — and a report that blends them cannot explain its own numbers.
The metrics that belong in the report
Report each engine separately. Profound published an answer-engine citation study in July 2025 finding that nearly 89% of AI citations come from completely different sources depending on which model users query. A blended score averages away exactly the differences the client needs to see; a per-engine scorecard shows where the brand is strong and where the work is.
Mentions and share of answer
Count how many tracked prompts mention the brand, per engine, and note where in the answer the mention sits — a first-sentence recommendation is not the same as a name in a closing list. Then compute share of answer, the brand's presence relative to every brand named across the tracked prompts; the method is in how to calculate share of answer. Record mention-without-citation cases separately: they show the engine knows the brand but is relying on third-party descriptions of it.
Citations and sources
For each engine, list the domains cited when the brand comes up, split between owned sources (the client's site) and third-party sources (press, directories, review platforms). Perplexity is the most explicit engine here: its developer documentation, verified on 8 September 2026, describes how the model inserts numbered references in the answer text, with the corresponding source URLs delivered alongside. That explicitness makes Perplexity citations the easiest per-engine KPI to defend in front of a client, since every count traces back to a visible source list.
Sentiment and accuracy
Whether the brand is mentioned is half the metric; how it is described is the other half. Add a sentiment line per engine — positive, neutral or negative framing — and an accuracy check on whether the engine describes the offer, pricing model and locations correctly. Consider a hypothetical answer that recommends a brand while attaching an outdated description of it: the mention would help less than it appears to.
Use this checklist to record each engine's answers. Check citation availability in each run; the table is a reporting worksheet, not a summary of Namedrop test results.
| Engine | Mentions | Citations |
|---|---|---|
| ChatGPT | Record whether and where the answer names the brand | Save any cited URLs; record their absence when none are provided |
| Perplexity | Record whether and where the answer names the brand | Match numbered references to source URLs, as described in Perplexity's documentation |
| Gemini | Record whether and where the answer names the brand | Save any cited URLs; record their absence when none are provided |
| Claude | Record whether and where the answer names the brand | Record whether web search was used and which URLs were cited |
The prompt set and the competitor benchmark
The trend line is only as good as the prompt set's stability. Agree a set of prompts with the client — the questions their buyers actually ask — and lock it. Changing prompts each cycle destroys the trend: month two no longer measures the same thing as month one, and every movement becomes unexplainable. When the business changes, add prompts with a marked start date and report them separately until they have a history.
Run each prompt several times per period, never once. OpenAI's documentation on the seed parameter, verified on 8 September 2026, states that determinism is not guaranteed and points to the system_fingerprint response parameter for monitoring backend changes. If the model's own vendor does not promise identical outputs for identical inputs, a single run is an anecdote; repeated runs across the month are a measurement.
Benchmark against a fixed competitor set on the same locked prompts, never on ad hoc queries. The client question is compared to whom, and on which prompts — so present the benchmark as a per-engine view of the brand against each named competitor on the shared prompt set, not as one abstract composite score. A score without named competitors and named prompts cannot be challenged, which also means it cannot be trusted.
For multi-location and local-service clients, build the set in two layers: a shared core of brand and category prompts, plus a per-location template — the same buyer question phrased for each city — instantiated per market. Aggregate results by location in the report, so a local manager sees their own line while the executive summary keeps the overall picture.
The report template, section by section
The skeleton below is the whole report, in reading order. The executive summary comes first, and every section after it should be answerable in one glance: the number, last period's number, why it moved. The example lines use Alpencab, an invented airport-transfer brand with realistic but fictional values.
Snapshot header. Client, period, engines covered, prompt-set version, run schedule. Example: Alpencab, August 2026, ChatGPT, Perplexity, Gemini and Claude, prompt set unchanged since March, every prompt run weekly.
Executive summary. Three or four plain sentences. Example: Alpencab is now mentioned in roughly half of tracked prompts on Perplexity, up from about a third last month; the gain follows the airport-guide pages published in July, which Perplexity now cites. ChatGPT still describes an outdated pricing model, so correcting the source pages is next month's priority.
Per-engine scorecard. One block per engine: mention rate, share of answer, owned versus third-party citations, sentiment — each beside last period's value. Semrush's reporting guide, consulted on 8 September 2026, arranges metrics in tiers based on how close they are to business value; apply the same logic and lead with what the client's management cares about.
Competitor benchmark. The fixed competitor set on the same prompts, per engine. Example: on category prompts, Gemini recommends TransferPro ahead of Alpencab in most runs, citing two review platforms where Alpencab is absent.
Cited sources. The domains each engine relied on this period, owned versus third-party, with new entrants flagged. Example: a regional travel blog entered Perplexity's citations this month, and its description of Alpencab is accurate.
Recommended actions. For every metric that moved, one what-changed-and-why line linking the movement to actions taken during the period, then the next actions with owners. This section answers the client's real question: which of the things they paid for worked.
When two tools disagree, and what to keep as evidence
Two tools tracking the same brand will report different numbers, and neither has to be wrong. They differ in prompt phrasings, run counts, run dates, model versions and, most often, in what counts as a mention — some count any appearance of the brand name, others only a recommendation. Reconcile in that order: compare definitions first, then the prompt sets, then run counts and dates. Most discrepancies dissolve at the definition step.
This is where the continuous monitoring under the monthly report earns its place. When the client questions a number, pull the individual runs behind it and show which answers, on which dates, produced the figure. A report that cannot be decomposed into raw answers is an assertion, not a measurement.
Keep the evidence deliberately. For every figure in the report, store the raw answer transcript, the date and time of the run, the engine, and the model version where the API exposes it. Then set a retention period and delete on schedule: transcripts can contain personal data, and CNIL's guidance on retention, verified on 8 September 2026, states that personal data cannot be kept indefinitely — the data controller must define a retention period tied to the purpose the data was collected for. A workable rule: keep transcripts while the figures they support sit in a live report, plus the contractual dispute window, then purge.
Where to start this week
Sequence matters more than speed here, because the one thing that cannot be fixed later is a prompt set that changed mid-engagement.
First, agree the prompt set with the client before any measurement: a couple of dozen prompts drawn from how their buyers actually ask, plus the per-location template if they operate in several markets. Get written sign-off — that sign-off protects the trend line later.
Second, run the multi-engine baseline with repeated runs spread across the week, and store everything: transcripts, dates, engines, model versions.
Third, fill the template once and label it the baseline. Month one establishes the numbers, month two establishes the trend — resist reading movement into a single period.
Fourth, schedule the monthly cadence and set up the underlying monitoring so runs accumulate between reports; how to automate AI visibility checks covers the setup.
Run a free Namedrop scan to generate the baseline numbers for your first client report.
Sources
- Pew Research Center, Google users are less likely to click on links when an AI summary appears in the results, consulted 2026-09-08
- Gartner, Inc., Gartner Predicts Search Engine Volume Will Drop 25% by 2026 Due to AI Chatbots and Other Virtual Agents, consulted 2026-09-08
- Profound, Answer Engine Citation Overlap Strategy: How to Win at AI Visibility, consulted 2026-09-08
- Perplexity, Streaming Citation Parsing - Perplexity, consulted 2026-09-08
- OpenAI, How to make your completions outputs consistent with the new seed parameter, consulted 2026-09-08
- Semrush, Create an AI Brand Visibility Report [+ Template], consulted 2026-09-08
- CNIL, Les durées de conservation des données, consulted 2026-09-08