To measure AI visibility across ChatGPT, Gemini, Claude, and Perplexity, run a stable set of buyer-relevant prompts on each platform, capture the complete answers, and track mentions, recommendations, citations, accuracy, and competitors separately.
The method matters more than the headline score. These products do not behave identically. Some answers use live web search and expose citations. Some do not. Interfaces, models, personalization, location, and time can all influence the result.
A credible report preserves that context and looks for repeatable patterns instead of pretending every response is a fixed search ranking.
Start with the decision you need to make
Do not begin by collecting prompts because they are easy to generate. Begin with the decision the data should support.
You might want to know where competitors dominate category shortlists, which product claims are misunderstood, which content earns citations, or whether a market launch is becoming visible. Each question produces a different prompt set and review cadence.
Write the objective at the top of the project. For example: Identify the most important content and source gaps preventing our analytics product from appearing in mid-market buyer shortlists in the United Kingdom.
Step 1: Build a balanced prompt library
Organize prompts by customer journey rather than keyword variation.
- Problem awareness: How can a team detect whether AI tools recommend competitors?
- Category education: What is AI visibility monitoring?
- Use case: Tools for a SaaS marketing team tracking brand citations in AI answers.
- Comparison: Compare two or three known approaches.
- Alternatives: What are alternatives to a market leader?
- Recommendation: Best platforms for a defined company size, region, or workflow.
- Objection: Is this category useful if the company already has an SEO platform?
Use language from sales calls, support tickets, site search, community discussions, and customer interviews. Keep prompts natural. A good starting baseline is thirty to fifty prompts for one product and market.
Version the library. Add or remove prompts deliberately so changes in the sample do not masquerade as changes in visibility.
Step 2: Define the test conditions
Record the platform, product surface, date and time, account state, market or locale, and whether web search was used. If the interface exposes a model name, record it, but expect product labels to change.
Do not combine a searched answer with a non-searched answer as if they were equivalent. ChatGPT can search the web and provide citations; Claude can use web search and cite sources; Perplexity is built around web-backed answers with source links; Gemini may expose sources or related links and offers a separate double-check feature. Availability and behavior can vary by product and account.
The safest reporting language is observational: In this sample, on these surfaces, during this period.
Step 3: Decide how often to run prompts
Weekly measurement works for active categories and campaigns. Monthly measurement can be enough for slower markets. Daily testing often creates more noise and cost than insight unless you are investigating a launch or a fast-moving event.
Use the same cadence across the baseline and comparison periods. If possible, distribute runs across more than one moment so one temporary response does not define the month.
Repeated sampling can improve confidence, but it also multiplies workload. State the run count clearly. Never hide one test behind a percentage that looks statistically strong.
Step 4: Capture the answer, not just the brand name
Store the complete response and the context needed to review it. Then extract structured fields:
- brand mentioned: yes or no;
- mention count, where useful;
- recommended or merely referenced;
- ordered position, only when the answer presents an order;
- competitors mentioned;
- owned-domain citation: yes or no;
- all exposed citation URLs and domains;
- description or claim about the brand;
- accuracy status and reviewer note;
- answer refusal, error, or no relevant result.
Keep the raw answer. Extraction rules will improve, and a reviewer may disagree with an automated label.
Step 5: Separate the metrics
Mention rate
Divide answers containing the brand by eligible answers in the segment. Report it by platform, topic, intent, and market. An overall number without segments can hide the most actionable gap.
Recommendation rate
Count only answers that present the brand as a suitable option. A sentence saying a product exists is not the same as a recommendation.
Citation rate
For answers where citations are available, track the share that link to your domain and the pages used. Keep the denominator platform-specific. A surface that did not expose citations should not be scored as a failed citation.
Competitive share of voice
Choose a documented formula. One simple version is brand mentions divided by all mentions among the tracked competitor set. Another is the share of eligible answers where each brand appears. Do not switch formulas between reports.
Message accuracy
Review material claims: category, audience, features, pricing status, availability, and limitations. Use a small rubric such as accurate, partly accurate, inaccurate, or unverifiable. High visibility with old positioning deserves attention.
Prompt coverage
Show which topics and journey stages contain any visibility. Coverage maps make it easier to see that a brand wins educational prompts but disappears from recommendations.
Step 6: Interpret each platform carefully
ChatGPT
Record whether Search was active and retain inline citations or the Sources panel when present. OpenAI says ChatGPT may automatically search when web information would help, and users can also choose Search. Do not assume every ChatGPT answer used the live web.
Gemini
Capture sources or related links when shown. Google notes that not every Gemini response includes them, and a link shown by the double-check feature is not necessarily a source used to generate the original answer. Preserve that distinction in your data.
Claude
Record whether web search was enabled. Anthropic says web-searched responses include citations and advises users to inspect the original sources because the synthesis may omit context or contain errors.
Perplexity
Capture the cited sources and search mode. Perplexity describes its product as web-searching and source-linked, but source presence does not make every claim correct. Review source quality, relevance, and whether the citation actually supports the statement.
Step 7: Turn patterns into actions
A dashboard is only useful if it changes the backlog.
- If competitors win comparison prompts, inspect their cited evidence and publish a more useful, honest comparison.
- If your brand is mentioned inaccurately, update the clearest owned facts and seek correction or corroboration where the old claim appears.
- If one strong page earns citations, connect it to relevant product and documentation pages and keep it current.
- If a third-party domain shapes many answers, evaluate whether legitimate editorial, partnership, directory, or community participation belongs in the plan.
- If all brands fluctuate, avoid reacting to one run. Extend the sample.
Our guide to mentions, citations, and recommendations explains why these signals should not be collapsed into one event.
A practical reporting template
Begin each report with scope: objective, prompt count, platforms, markets, dates, run count, and material limitations. Then show:
- overall and platform-level mention and recommendation rates;
- competitive share by journey stage;
- citation domains and top cited pages;
- message accuracy issues;
- the five largest repeated gaps;
- recommended owners and next actions;
- changes from the comparable prior period.
Link every aggregate back to the underlying answers. A stakeholder should be able to inspect why a response was classified.
Common measurement mistakes
- Changing prompts every month without preserving a baseline cohort.
- Testing only the company name, which measures recognition rather than discovery.
- Treating all mentions as positive recommendations.
- Penalizing a platform for missing citations when that answer surface did not expose any.
- Ignoring locale, search state, and account context.
- Calling an observed score an official platform metric.
- Making causal claims from a before-and-after change with no control or supporting evidence.
How Swep helps
Swep is designed to make this operating rhythm easier: manage prompt sets, monitor supported AI surfaces, identify brands and competitors in answers, inspect citations, and organize opportunities. The aim is consistent observation and better prioritization, not a promise to control independent models.
You can begin with the manual framework in this article, then use the comparison page and documentation to decide whether a dedicated workflow is justified.
The takeaway
Measure AI visibility like a research program, not a vanity leaderboard. Fix the prompt set, preserve platform context, retain the raw answers, separate the metrics, and be honest about uncertainty.
The best report does not say that a brand is “72% optimized for AI.” It shows exactly where the brand is included, excluded, cited, misunderstood, or beaten—and what the team should do next.
