Our methodology
Every WhoCitesYou audit runs the same way, so scores are comparable and defensible. We run the buyer-intent questions in your category across three AI engines, repeatedly and in fresh chats, capture every answer verbatim, classify each mention, and score the result 0 to 100 with an automated scoring engine.
From query matrix to scored report
We build your buyer-intent query matrix
We map the questions a real buyer would ask an AI in your category: 12 queries for a full audit, 3 for a free mini-audit, split across category, comparison, alternative, and use-case intent. These are the exact moments where an AI hands out recommendations, so the matrix targets where citations are actually won and lost.
We run every query across three engines, in repeated trials
Each query is run through ChatGPT, Gemini, and Perplexity, with repeated trials per query rather than a single ask. For a full audit that is 12 queries × 3 engines × repeated trials; a mini-audit samples 3 queries × 3 engines × 2 trials for 18 captures. AI answers behave like distributions, so we sample them like distributions.
Every query runs in a fresh chat
No query reuses a session. Each one starts clean, with no prior turns, no account history, and no memory carried over, so one answer cannot prime the next. This is what keeps captures comparable across brands and engines instead of contaminated by conversation context.
Answers are captured verbatim
We record what the engine actually said, word for word, not a paraphrase or a remembered gist. The verbatim text is the evidence: every classification, quote, and score traces back to a real captured answer you could reproduce yourself.
Each mention is annotated
For every capture we record whether the brand was cited, invisible, or misrepresented, plus the mention strength (from a passing name-drop to the outright top recommendation) and the first-mention rank (how early the brand appears relative to competitors). Competitors named instead are logged in the same pass.
Share of voice is computed against competitors
Across every captured answer, we tally how often each brand is named to produce a share of voice: who wins your category in AI answers, by how much, and which of your queries are being handed to rivals instead.
An automated scoring engine produces the 0-100 score
Our software applies the same rubric to every audit: mention rate, mention strength, first-mention rank, and accuracy roll up into a single 0-100 visibility score, broken out per engine. Because the scoring is automated and identical across clients, two audits are directly comparable.
Cited, invisible, or misrepresented
Most people imagine AI visibility as a yes-or-no question: were you mentioned or not? We track a third state, because the most damaging outcome is not silence.
The engine names you
Your brand appears in the answer. We still record how strongly and how early, because being listed fifth as an afterthought is not the same as being the first recommendation.
The engine skips you
The answer names competitors and moves on without you. This is capturable demand handed straight to rivals, and it is invisible to you unless you go looking for it.
The engine gets you wrong
You are named, but the engine states something inaccurate: a wrong price, an out-of-date feature, a stale positioning. A confident, wrong answer can be worse than no mention at all.
Three decisions that make the numbers trustworthy
Each choice in the protocol exists to remove a specific way an AI-visibility read can lie to you.
To avoid personalization contamination
An assistant that remembers your last few questions will skew toward them. Reusing a session lets one answer prime the next and quietly bends the result. A fresh chat per query removes that contamination so the capture reflects what a new buyer would actually see.
To measure run-to-run variance
Ask the same question twice and you can get a different shortlist, in a different order. A single screenshot proves nothing. Repeated trials turn a one-off answer into a rate, so your citation rate and share of voice reflect what typically happens, not one lucky or unlucky roll.
Because a confident wrong answer is worse than silence
An engine that quotes a wrong price or a stale feature actively misinforms a buyer at the moment of decision. Folding that into a plain mention would hide real damage, so misrepresentation is scored as its own, separate failure state.
Automated software, human-checked before delivery
The audit is a software pipeline: it runs the query matrix, captures answers, applies the classification and scoring rubric, and generates a self-contained HTML report. It is a productized, automated tool, not a consulting engagement, and every report is checked before it ships. Because AI answers vary between runs, every audit is a defensible sample of engine behavior, not a permanent census.
The same methodology, on four real brands
Every sample below was produced with the exact protocol on this page. Read them in full: two mini-audits, two full audits, one winning AI search and one losing it.
Tabby
UAE buy-now-pay-later. Cited in all 18 sampled answers, first pick on the flagship query across every engine.
Read the report →Cal.com
Scheduling software. Leads share of voice, yet the engines still frame it as the Calendly alternative.
Read the report →beehiiv
Newsletter platform. Strong visibility with a 100% ChatGPT mention rate, now defending the lead.
Read the report →Ziina
UAE fintech. Invisible on 10 high-intent answers, out-positioned by rivals, with one factual misrepresentation.
Read the report →Run this methodology on your brand: free mini-audit
Tell us your website and we'll run a free 1-page sample with the same protocol: 3 real buying queries across ChatGPT, Gemini, and Perplexity, plus one clear insight on where you stand. The full 12-query AI Visibility Audit is $197 during founding-client pricing. To see this protocol applied across a whole sector, read our UAE fintech AI-visibility roundup.