Home / Methodology
How the audit works

Our methodology

Every WhoCitesYou audit runs the same way, so scores are comparable and defensible. We run the buyer-intent questions in your category across three AI engines, repeatedly and in fresh chats, capture every answer verbatim, classify each mention, and score the result 0 to 100 with an automated scoring engine.

12 full / 3 mini queries 3 engines Repeated trials Fresh chat per query 0-100 score
The protocol

From query matrix to scored report

STEP 01

We build your buyer-intent query matrix

We map the questions a real buyer would ask an AI in your category: 12 queries for a full audit, 3 for a free mini-audit, split across category, comparison, alternative, and use-case intent. These are the exact moments where an AI hands out recommendations, so the matrix targets where citations are actually won and lost.

STEP 02

We run every query across three engines, in repeated trials

Each query is run through ChatGPT, Gemini, and Perplexity, with repeated trials per query rather than a single ask. For a full audit that is 12 queries × 3 engines × repeated trials; a mini-audit samples 3 queries × 3 engines × 2 trials for 18 captures. AI answers behave like distributions, so we sample them like distributions.

STEP 03

Every query runs in a fresh chat

No query reuses a session. Each one starts clean, with no prior turns, no account history, and no memory carried over, so one answer cannot prime the next. This is what keeps captures comparable across brands and engines instead of contaminated by conversation context.

STEP 04

Answers are captured verbatim

We record what the engine actually said, word for word, not a paraphrase or a remembered gist. The verbatim text is the evidence: every classification, quote, and score traces back to a real captured answer you could reproduce yourself.

STEP 05

Each mention is annotated

For every capture we record whether the brand was cited, invisible, or misrepresented, plus the mention strength (from a passing name-drop to the outright top recommendation) and the first-mention rank (how early the brand appears relative to competitors). Competitors named instead are logged in the same pass.

STEP 06

Share of voice is computed against competitors

Across every captured answer, we tally how often each brand is named to produce a share of voice: who wins your category in AI answers, by how much, and which of your queries are being handed to rivals instead.

STEP 07

An automated scoring engine produces the 0-100 score

Our software applies the same rubric to every audit: mention rate, mention strength, first-mention rank, and accuracy roll up into a single 0-100 visibility score, broken out per engine. Because the scoring is automated and identical across clients, two audits are directly comparable.

Three outcomes, not two

Cited, invisible, or misrepresented

Most people imagine AI visibility as a yes-or-no question: were you mentioned or not? We track a third state, because the most damaging outcome is not silence.

Cited

The engine names you

Your brand appears in the answer. We still record how strongly and how early, because being listed fifth as an afterthought is not the same as being the first recommendation.

Invisible

The engine skips you

The answer names competitors and moves on without you. This is capturable demand handed straight to rivals, and it is invisible to you unless you go looking for it.

Misrepresented

The engine gets you wrong

You are named, but the engine states something inaccurate: a wrong price, an out-of-date feature, a stale positioning. A confident, wrong answer can be worse than no mention at all.

Why the method is built this way

Three decisions that make the numbers trustworthy

Each choice in the protocol exists to remove a specific way an AI-visibility read can lie to you.

Why fresh chats

To avoid personalization contamination

An assistant that remembers your last few questions will skew toward them. Reusing a session lets one answer prime the next and quietly bends the result. A fresh chat per query removes that contamination so the capture reflects what a new buyer would actually see.

Why repeated trials

To measure run-to-run variance

Ask the same question twice and you can get a different shortlist, in a different order. A single screenshot proves nothing. Repeated trials turn a one-off answer into a rate, so your citation rate and share of voice reflect what typically happens, not one lucky or unlucky roll.

Why a misrepresented category

Because a confident wrong answer is worse than silence

An engine that quotes a wrong price or a stale feature actively misinforms a buyer at the moment of decision. Folding that into a plain mention would hide real damage, so misrepresentation is scored as its own, separate failure state.

Automated software, human-checked before delivery

The audit is a software pipeline: it runs the query matrix, captures answers, applies the classification and scoring rubric, and generates a self-contained HTML report. It is a productized, automated tool, not a consulting engagement, and every report is checked before it ships. Because AI answers vary between runs, every audit is a defensible sample of engine behavior, not a permanent census.

Start here

Run this methodology on your brand: free mini-audit

Tell us your website and we'll run a free 1-page sample with the same protocol: 3 real buying queries across ChatGPT, Gemini, and Perplexity, plus one clear insight on where you stand. The full 12-query AI Visibility Audit is $197 during founding-client pricing. To see this protocol applied across a whole sector, read our UAE fintech AI-visibility roundup.