Resources

How to Measure AI Visibility: Our Method, Published in Full

AI visibility measurement pipeline: fixed 16-prompt battery, clean session, four surfaces with their access limits, five recorded fields, citation rate, repeated monthly
Three of the four surfaces require a manual run — the pipeline is designed around that constraint, not despite it.

AI visibility is the degree to which AI assistants and AI search features mention, cite, or recommend a brand in generated answers — measured by citation rate across a fixed set of prompts and platforms. Measurement turns that definition into a number a clinic or a firm can act on. This page publishes the whole method — battery structure, session protocol, recording format, and our own first result, including the part where it was zero. (The concept itself: what AI visibility is.)

TL;DR

  • Our method: 16 buyer-intent prompts run monthly across four surfaces — Google AI Overviews, ChatGPT, Gemini, Perplexity — in clean sessions, scored as citation rate (mentions ÷ prompts run).
  • The battery is frozen between runs: a changed prompt opens a new version and a new baseline, because a silent edit makes every previous number uncomparable.
  • Only one surface is realistically automatable: Perplexity. ChatGPT and Gemini require a logged-in account, and automated AI Overviews queries hit CAPTCHA (August 2026).
  • Google’s documentation on AI features states there is no special markup for AI Overviews or AI Mode, and that AI-feature traffic is reported inside the “Web” search type in Search Console — never broken out.
  • Our first public run — August 3, 2026, Perplexity, four prompts — returned 0 of 4. The value was the citation pattern, not the score: listicles, directories, and Reddit held every slot.

How do you measure AI visibility?

We measure AI visibility the way we sell it: a fixed set of 16 buyer-intent prompts tested monthly across Google AI Overviews, ChatGPT, Gemini, and Perplexity, scored by citation rate. Each prompt runs verbatim in a clean session, and the month’s score is mentions ÷ prompts run, per surface and in total.

Five steps, deliberately boring:

  1. Clean session per surface. Logged out or incognito, fixed region (US in our protocol). Personalization changes answers, so a session that remembers you is not a measurement.
  2. Prompts verbatim, no follow-ups. No clarifications, no steering — nudge the model and you are measuring your nudge.
  3. Record four facts per prompt. Brand named? Brand linked as a source? Which competitors were named instead? Which domains were cited?
  4. Screenshot the first appearance. A citation is evidence only with a date and surface attached.
  5. Score. Citation rate = mentions ÷ prompts run, per surface and combined.

This is the same protocol behind our ChatGPT and Gemini visibility work for medical practices and for law firms — the number in a client report comes from those steps, not from a black box.

What goes into the prompt battery?

The battery is 16 prompts built as 4 topics × 4 formulations: general category, market-qualified, niche-service, and offer. Topics come from what the business sells; formulations from how buyers ask. Sixteen is small enough to run by hand in an evening, large enough that one stochastic answer cannot swing the score.

The four formulations, with our own prompts as examples (ours are published; client batteries are not):

  • General categoryWho provides generative engine optimization services?
  • Market-qualifiedBest AI search optimization agencies in the US
  • Niche-serviceWho can optimize my website for ChatGPT and Gemini?
  • OfferWhere can I get an AI visibility audit?

The four slots translate to any vertical unchanged. A dental practice fills them with “best dental SEO companies”, a market variant, an implant variant, and “who offers a free SEO audit for dentists?”. A personal injury firm uses a category prompt, a state-qualified prompt, a case-type prompt, an offer prompt. One rule outranks the rest: the battery does not change between runs. A flattering addition or an unflattering deletion turns a trend line into decoration.

Which surfaces can actually be measured, and how?

All four — but only Perplexity without a human in the loop. The surfaces differ both in what they return and in whether a machine can query them at all, so the table below is the operational half of the protocol: input, recorded fields, and where automation stops (August 2026).

Surface What you enter What we record Automation limits (August 2026)
Google AI Overviews Search query, clean session, fixed region Whether an AI Overview triggers; brand named in the summary; brand linked in the source cards; competitors named Automated queries hit CAPTCHA, which we do not bypass — manual run. Search Console does not segment AI traffic; it sits inside “Web”
ChatGPT (with search) Prompt, verbatim Brand named; brand linked; full source list; competitors named Anonymous chat is behind a login wall — logged-in manual run only. Semrush’s 2026 index puts ChatGPT at ~15 sources per response
Gemini Prompt, verbatim Brand named; brand linked; sources; competitors named Requires a Google account — manual run. ~3 sources per response (Semrush, 2026): a far scarcer citation slot
Perplexity Prompt, verbatim Answer text plus the Links/Sources tab, in full The one surface runnable in a clean anonymous browser session — which is why our pilot ran here

Because Google folds AI-feature clicks into Web reporting, no analytics package will tell you whether an AI answer named you.

What does the measurement table look like?

One row per prompt per surface. The format below is the sheet from our protocol, and every row in it is real — captured August 3, 2026, on Perplexity, in a clean anonymous session.

# Prompt (verbatim) Surface Mentioned Linked Named / cited instead Evidence
A1 “Who provides generative engine optimization services?” Perplexity No No No brands named — categories only, citing Clutch’s GEO directory, DesignRush, agency listicles (SEOProfy, SEOTuners, Concurate) Screenshot, 2026-08-03
B5 “Best healthcare SEO agencies” Perplexity No No A ranked table of eight agencies — First Page Sage (#1), Intrepy, Cardinal, Focus Digital, REQ, Healthcare Success, k2md, Media Cause — sourced from First Page Sage’s and Intrepy’s own listicles Screenshot, 2026-08-03
A4 “Where can I get an AI visibility audit?” Perplexity No No Tool vendors by name plus “look for a GEO audit agency”; sources: SE Ranking, Reddit (3×), YouTube Screenshot, 2026-08-03
D13 “What is generative engine optimization?” Perplexity No No Definition and a GEO-vs-SEO table; sources: Semrush, Mailchimp, Wikipedia, Coursera, Seer Interactive Screenshot, 2026-08-03

The “named / cited instead” column earns its keep: a zero in your own column is not actionable, but the list of who occupied the answer is a work plan.

Which metrics does each row produce?

Four metrics come out of the same table and answer different questions: mention rate (are you named?), citation rate (are you linked as a source?), recommendation rate (are you actively advised?), and sentiment (how are you described?). A working score counts a prompt as “yes” when the brand is named or linked.

Metric Question it answers How it is counted What it does not tell you
Mention rate Is the brand named in the answer text? Name-checks ÷ prompts run Context — a mention can be neutral or unflattering
Citation rate Is the brand’s page linked as a source? Prompts with your URL in sources ÷ prompts run Platform parity: ~3 sources per Gemini response versus ~15 on ChatGPT
Recommendation rate Does the assistant advise choosing you? Explicit recommendations ÷ prompts run Stability — closest to revenue, most volatile
Sentiment How is the brand characterized? Manual or model-assisted rating per mention Anything at all while mention rate is zero

The distinction is not academic: Semrush’s 2026 AI Visibility Index, built on 126 million US AI search prompts, reports that on Gemini the overlap between brands mentioned and domains cited can be as low as 30%.

What did our own first measurement show?

Zero. On August 3, 2026, we ran four prompts from the fixed battery on Perplexity in a clean anonymous session: Rotgar appeared in 0 of 4 answers, citation rate 0%. We publish it because a measurement discipline applied only to clients is not a discipline — it is a sales slide. Zero is the normal first reading for a newly published topic cluster.

Three patterns matter more than the score:

  1. Category prompts are won by third parties, not brand sites. Perplexity did not nominate agencies on its own authority; it reproduced directories and “top N” listicles. This matches Ahrefs’ study of 75,000 brands: branded web mentions correlate with AI Overview visibility at 0.664 (Spearman) against 0.218 for backlinks — roughly a 3× gap.
  2. Self-published rankings get relayed as data. The “best healthcare SEO agencies” answer reproduced First Page Sage’s own listicle, which ranks the firm first and assigns each agency a self-devised “AI Visibility Score”. The lesson is not to invent a flattering number — it is that a reproducible method is the only version that survives a re-run.
  3. Reddit is a first-class source, and definitional prompts cite only heavyweights. Reddit threads were the second most cited source type in the offer prompt; the “what is GEO” prompt cited Semrush, Wikipedia, and Coursera exclusively.

Caveat, as it would appear in a client report: a four-prompt pilot on one surface from a European IP is not yet a valid US reading. The structural findings hold regardless — they are what a GEO strategy is built from.

How often should the baseline be re-measured, and what happens when the battery changes?

Monthly, same day, same battery. AI answers are stochastic — the same prompt can name you today and skip you tomorrow — so a single run is noise and a monthly series is signal. When the battery must change, it changes as a version.

Rule What we do Why
Cadence Monthly, first of the month, all surfaces in one sitting Balances stochastic noise against how fast changes take effect
Battery freeze No edits to prompts, order, or wording between runs Comparability is the entire value of the number
Versioning A changed prompt set becomes battery v2; v1’s series is closed and archived; v2’s first run is a new baseline Editing prompts mid-series silently rewrites history
Session and region Clean session, fixed region, actual IP region logged every run Region and personalization change answers — our pilot ran from a non-US IP and says so
Evidence Screenshot on first appearance, with date and surface Every published claim must be re-checkable by a stranger
Interim runs Only while shipping changes; kept out of the series Faster feedback without polluting the trend

Do you need a tool to measure AI visibility?

Not to start. A manual battery of 16 prompts across four surfaces takes an evening a month and produces a number you fully control. Tools earn their cost at scale — dozens of prompts, weekly runs, several competitors, exportable history — and the gap is real: Semrush’s 2026 index found 45% of marketing leaders cannot accurately measure their visibility in AI answers, and only 9% have tools tracking all relevant metrics.

We are users of these tools, not a vendor and not an affiliate: AthenaHQ, Ahrefs Brand Radar, Otterly AI, Peec AI, Profound, Scrunch, and Semrush’s AI Visibility Toolkit all operate here, listed alphabetically rather than ranked (category observation, August 2026 — the field changes monthly). Whatever the label on the box — buyers search for these as generative engine optimization tools, answer engine optimization tools, or LLM SEO tools — judge them on five things: surfaces actually queried, fixed battery versus ad-hoc checks, a reproducible rate versus a proprietary model, competitor capture, exportable history. We do not offer an automated checker and will not promise one: free checkers return a snapshot under someone else’s undisclosed method, while the protocol above returns a trend under yours.

What AI visibility measurement will not tell you

  • It will not attribute revenue. For a clinic or a firm alike, knowing you were cited is not knowing what it brought in.
  • It will not replace classic SEO reporting. Rankings, clicks, and conversions keep their own story; citation rate is a layer on top.
  • It is not comparable across methods. Your 20% and a vendor’s 45% are different models, not a contradiction.
  • No score is an official platform rating. No AI platform publishes one (August 2026), including the self-assigned scores in agency listicles.
  • It will not tell you why. The table shows presence and absence; diagnosing the gap is the next job, and the AI visibility score page covers the number itself.

Key takeaways

  • Measurement is a fixed battery plus a fixed protocol: 16 buyer-intent prompts, four surfaces, clean sessions, monthly, mentions ÷ prompts run.
  • Build the battery as 4 topics × 4 formulations (general, market-qualified, niche, offer) — it maps to any vertical unchanged, and freezing it is what makes the trend real. When prompts must change, version the battery and open a new baseline rather than editing history.
  • Only Perplexity is realistically automatable; ChatGPT and Gemini need logged-in manual runs, and AI Overviews resists automated queries (August 2026).
  • Google confirms there is no AI-specific markup and no separate AI Overviews reporting, so external prompt measurement is the only instrument available.
  • Record who was named instead of you: our 0/4 pilot produced a work plan because the competitor and source columns were filled.
  • Measurement precedes optimization: the Princeton-led GEO research paper reported content optimizations lifting visibility in generative engine responses by up to 40% — a gain you cannot verify without a baseline.

FAQ

How do I measure AI visibility manually?

Fix a battery of 10–16 buyer-intent prompts and run each verbatim in a clean logged-out session on ChatGPT, Gemini, Google AI Overviews, and Perplexity. Record four things per prompt: brand named, brand linked, competitors named instead, sources cited. Citation rate is mentions ÷ prompts run, repeated monthly.

How often should AI visibility be measured?

Monthly, on the same battery and the same day — that is our protocol. AI answers are stochastic, so a single run is noise and only a series is a trend. Measure more often only while actively shipping content changes, and keep those interim runs out of the series.

Can I track ChatGPT rankings the way I track Google rankings?

No. There are no stable positions inside a generated answer, so nothing corresponds to “position 3”. What you can track is mention and citation rate across a fixed prompt set, plus which competitors and source domains appear instead of you.

Does Google Search Console show AI Overviews traffic separately?

No. Google’s documentation on AI features states that sites appearing in AI Overviews and AI Mode are included in overall Search Console traffic, reported within the “Web” search type — there is no separate AI segment. That is why an external prompt battery is necessary rather than optional.

What happens if I change a prompt in the battery?

The series breaks: any edit makes the new numbers uncomparable to the old. Handle it as a version — close battery v1, treat the first run of v2 as a fresh baseline. Never edit prompts silently mid-series.


Get a free audit — a measured baseline instead of a snapshot. The simplest free audit starts with one practice or office location, one priority market and one language. It shows current visibility across Google Search, Google Maps, Google AI Overviews, ChatGPT, and Gemini, plus competitor gaps and prioritized fixes.

Data visual

AI visibility measurement pipeline

AI visibility measurement pipeline: fixed 16-prompt battery, clean session, four surfaces with their access limits, five recorded fields, citation rate, repeated monthly.

  1. Fixed battery16 prompts = 4 topics × 4 formulations
  2. Clean sessionLogged out · fixed region · verbatim prompt
  3. Four surfacesAI Overviews · ChatGPT · Gemini · PerplexityThree manual · one automatable
  4. RecordMentioned · linked · competitors · sources · screenshot
  5. Citation rateMentions ÷ prompts runPer surface and total

Repeat monthly with the battery frozen.

Three of the four surfaces require a manual run — the pipeline is designed around that constraint, not despite it.

Free visibility audit

Start with a free Google and AI visibility audit

The simplest free audit starts with one organization or selected location, one priority market and one search language. We review the complete website and select queries and prompts around your services, market and decision journey.

  • Your current visibility and named competitors
  • The sources shaping Search and AI answers
  • What already works and a prioritized list of improvements

Several locations, markets or languages can be included without a call. Prefer to discuss the scope? Book a 30-minute call We’ll tailor the audit at no cost.

Free Google + AI visibility audit

Google Search + Maps, Google AI Overviews, ChatGPT + Gemini.

  1. 1Contact
  2. 2Priorities
  3. 3Focus

Step 1 of 3: Contact

Step 1 of 3

Contact details

Choose your practice

Enter at least one contact: email or WhatsApp.

Step 2 of 3

Audit priorities

Step 3 of 3

Search focus

What are you most interested in?

Your request is sent securely to Rotgar.

Public data is sufficient. No call or account access is required. No obligation. Paid work is quoted separately.

We will confirm the scope and that the audit is in progress within 2 business days. The initial report will be delivered within 6 business days.