·
Skip to main content

The GEO journal8 min read

How to Do Prompt Research to Find AI Visibility Gaps

Prompt research is the practice of finding the real questions users ask ChatGPT, Perplexity and Gemini, then checking whether your brand appears in the answers. A keyword list cannot substitute for it, because prompts carry context keywords strip away. This workflow ends in a concrete gap list and runs without an enterprise platform.

What prompt research is, and what it is not

Two practices that share a name

The phrase covers two unrelated practices. The first is building libraries of prompts that make an AI assistant perform SEO tasks: drafting title tags, clustering keywords, summarising briefs. Useful, but not this article. The second is researching the prompts real users type into ChatGPT, Perplexity and Gemini, then testing whether your brand shows up in the answers those prompts produce. Everything below covers the second practice.

The distinction matters because AI-driven search changes how people discover and evaluate brands. Short keyword strings give way to longer, conversational questions, often refined over several turns of the same conversation. The engine replies with a synthesised answer and a short list of cited sources. Whether you appear in that answer is the new visibility question, and it cannot be read off a rank tracker.

How a prompt differs from a keyword

A keyword compresses intent; a prompt states it. Someone who types "project management software" could be a student, a buyer or a journalist. Someone who asks "which project management tool suits a ten-person design agency that bills by the hour" has told the engine who they are, what they need and what constraint matters.

Engines also process prompts differently. Google's Search Central documentation, verified on 22 September 2026, states that both AI Overviews and AI Mode "may use a 'query fan-out' technique — issuing multiple related searches across subtopics and data sources — to develop a response." One prompt can trigger many sub-searches, so a page can enter an answer through a sub-query nobody typed.

None of this makes keyword research obsolete. It remains the seed layer: the personas, intent classifications and long-tail queries you already maintain are the raw material prompt research starts from. Prompt research is what keyword research evolves into for AI search, not a replacement built from scratch.

DimensionKeyword researchPrompt research
Unit of researchShort query stringFull conversational question, often multi-turn
Intent visibilityInferred from the keywordStated inside the prompt
Data sourcesVolume databases, Search ConsoleYour own question data, engine suggestions, tracking tools
Volume dataAvailable in tool databasesLargely unavailable; prompts behave like extreme long tail
Output artifactPrioritised keyword listPrompt set with per-engine visibility status
Success metricRankings and clicksMentions and citations in answers

The table's last row is the operational shift: rankings are ordinal and public, while mentions and citations are binary per run and only visible if you go and check. That checking step is the workflow the rest of this article describes.

Where real prompt data comes from without an enterprise platform

Long-tail question queries are the bridge between the keyword list you have and the prompt list you need. Three sources supply them without a platform subscription.

Your own data: Search Console, site search, support tickets

The Performance report in Google Search Console supports regular expression filters on queries. Google's help documentation, verified on 22 September 2026, confirms that "regular expression (regex) search enables you to filter for, or exclude, multiple queries or URLs". A filter such as ^(how|what|which|why|can|should|is|does) isolates the question-style queries that already bring you impressions; the longest of them read like prompts. Add your site-search logs and, above all, the questions your sales and support colleagues answer every week — those are prompts users speak out loud. Support tickets deserve their own pass: the way a frustrated customer phrases a problem is usually closer to a real prompt than anything a keyword tool suggests.

The engines themselves: follow-ups, related questions, autocomplete

The engines expose real phrasing themselves. Ask a seed question in ChatGPT, Perplexity or Gemini and note the follow-up questions the interface suggests beneath the answer: they show how the conversation actually continues. Google's autocomplete on question stems ("how do agencies…", "is it worth…") does the same for the entry stage. Collect these verbatim rather than rewriting them; the phrasing is the data.

Where paid tools fit, and when you need one

Paid tools enter as a scale option, not a prerequisite. Semrush's documentation for its AI visibility features, verified on 22 September 2026, states that "based on your domain and location, Semrush generates a set of synthetic prompts related to your business (a mix of branded and non-branded queries)". Note the word synthetic: these prompts are generated, not observed, so treat them as an expansion source to review, not as user data. The honest threshold for buying tooling is operational: when the prompt set must be re-run on a schedule across several engines, manual checking stops scaling and automation becomes the cheaper path.

From seed prompts to a visibility gap list in five steps

Step 1: build the seed set from personas and keywords

Start from the personas and intent classifications your keyword research already uses. For each persona, write the questions they would ask at three journey stages: entry (understanding the problem), comparative (weighing options) and decision (choosing one). Decision-stage prompts get priority on a commercial site: they name product categories, and absence there costs the most.

Step 2: expand each seed into variants and follow-ups

Expand each seed the way you expand a head keyword: vary the phrasing, add constraints (budget, team size, industry, region) and append the follow-ups you collected from the engines. Then lock the set. A locked set is what makes later runs comparable; edits mid-stream turn your data into anecdotes.

Step 3: run the set against ChatGPT, Perplexity and Gemini

Run every prompt in a fresh conversation on each engine, signed out or with personalisation off where the product allows it, and record the date of every run. Generative answers vary between runs of the same prompt, so a single run is a snapshot, not a measurement; treat any one answer as one data point and plan to repeat the run before drawing conclusions.

By hand, the check is the same on each engine: paste the prompt, read the full answer, then inspect the sources the interface attaches — expand any collapsed source list before recording. Search the answer text for your brand name and for each competitor on your list; search the linked domains the same way. Ten prompts across three engines is roughly an hour of careful work.

Step 4: record mentions and citations per prompt

Record two outcomes separately: a mention, where the answer names your brand in its text, and a citation, where the answer links to your site as a source. The two come from different mechanisms and diverge by engine, which is why the deliverable tracks both — the distinction is unpacked in the difference between an AI mention and an AI citation.

Step 5: keep the prompts where competitors appear and you don't

Filter the recorded runs down to prompts where a competitor is mentioned or cited and you are absent. That subset, sorted by journey stage and priority, is the visibility gap list — the deliverable the next section fills in. Do not discard the prompts where you already appear: they become your regression set, the ones you re-check to make sure new content does not trade one gap for another.

What a finished prompt research deliverable looks like

The stakes justify the spreadsheet. Pew Research Center's study of Google users, published on 22 July 2025, found that users who encountered an AI summary clicked a traditional search result link in 8% of visits, against 15% when no summary appeared — nearly twice as often. Where the answer layer absorbs the click, presence inside the answer is the visibility.

The deliverable needs seven columns: prompt, journey stage, engine, date of run, your status (cited, mentioned or absent), competitors present, and priority. The table below is a filled-in hypothetical example — Plansmith, TaskRail and Boardline are invented brands, and every value is illustrative, not observed.

PromptStageEngineRun datePlansmith statusCompetitors presentPriority
best project tracker for a small design agencyDecisionChatGPT2026-09-18AbsentTaskRail, BoardlineHigh
how do agencies collect client feedback on deliverablesEntryPerplexity2026-09-18Mentioned, not citedTaskRailMedium
Plansmith vs TaskRail for freelancersComparativeGemini2026-09-19CitedTaskRailLow

Priority combines journey stage with competitive pressure: a decision-stage prompt where several competitors appear and you do not outranks an entry-stage prompt where nobody is cited yet.

Each row maps to a content decision. The first row — absent on a decision-stage prompt where two competitors appear — calls for creating a page that answers that exact question, typically a comparison or best-for page. The second — mentioned but not cited — calls for rewriting an existing page so it becomes citable. The third is working; leave it alone and re-check on the next run. The full method for that translation is covered in turning AI search data into a content gap list.

Where to start this week

A first pass fits in an afternoon.

  1. Pull question-style queries from the Search Console Performance report with the regex filter above.
  2. Add the questions your sales and support colleagues hear most often.
  3. Write a couple of dozen prompts covering all three journey stages, with decision-stage prompts first.
  4. Run each prompt once against two engines, in fresh conversations.
  5. Record the outcome per prompt — cited, mentioned or absent — with the date of the run, in the deliverable format above.

Scope this honestly. One run produces a snapshot, not a trend, and the filled-in table is an internal artifact, not a report to publish. Repeat the run on the same locked set before treating any gap as confirmed. Automation is the step after the manual pass proves the prompt set is worth tracking — the options are laid out in automating AI visibility checks once the prompt set is stable.

To see which prompts already surface your brand before you build the full set, run a free Namedrop scan as your baseline.

Sources

Frequently asked questions

Does keyword research still matter for AI search, or is it obsolete?
It still matters, as the seed layer. The personas, intent classifications and long-tail question queries you already maintain are the input prompt research starts from. What changes is the output: instead of a ranked keyword list, you produce a set of conversational prompts and a per-engine record of whether your brand is mentioned or cited. Treat prompt research as the extension of keyword research, not its replacement.
How do I identify content gaps using AI search data?
Start from the visibility gap list: the prompts where an engine mentions or cites competitors but not you. Each gap maps to a decision — create a page that answers the prompt directly, rewrite an existing page so it becomes citable, or leave a working page alone. Group gaps by journey stage and handle decision-stage prompts first, since they sit closest to purchase.
What tools track brand mentions across ChatGPT, Gemini, and Perplexity?
Three categories exist. Manual checking works at small scale: run each prompt by hand and record the result with a date. DIY automation uses the engines' APIs and a script to repeat runs on a schedule. Monitoring tools, Namedrop among them, run a prompt set across engines and report mentions and citations over time. Choose by scale: move up a category when re-running the set by hand stops being sustainable.
How many prompts do I need before the results mean anything?
Enough to cover each persona and journey stage — in practice a few dozen — run more than once. Generative answers vary between runs of the same prompt, so a single pass over a small set is a snapshot, not a trend. Treat the first pass as a way to test whether your prompt set asks the right questions, and only draw conclusions from gaps that persist across repeated, dated runs.