Methodology

How I measure AI visibility — and what the numbers can’t prove.

Most providers in this category report a score without saying which engines they asked, on what date, from which country, or how many times. That is not a measurement — it is a claim. This page documents the method behind every audit I deliver, including its limits.

01What gets measured

Five questions, not one score.

MentionDoes the engine name you at all?

The most basic question, and the one most businesses fail. Being absent from the answer is a different problem from being described badly, and it has a different fix.

CitationDoes it link to you as a source?

Distinct from a mention. An engine can describe your category accurately while citing a competitor's page as the source. Copilot is currently the only engine that reports this back to publishers first-party.

RecommendationAre you named as a suggested option?

The commercially meaningful one. Appearing in a list of examples is not the same as being the answer to “who should I hire.”

AccuracyIs what it says about you correct?

An inaccurate mention can be worse than none — wrong service, wrong location, wrong pricing, or confusion with a similarly-named business.

Share of voiceWho gets named instead of you?

Measured on the same question, at the same time, against named competitors — because “you're not very visible” is not actionable, and “on this query they are named and you are not” is.

02How every observation is recorded

A result without conditions attached is noise.

AI answers vary by session, by how the question is phrased, and by the country the query comes from. The same prompt can name you at 9am and omit you at 3pm. So every observation carries the conditions it was captured under:

  • The exact prompt— fixed in advance, not adjusted afterwards to produce a better-looking result.
  • The engine— reported separately, never merged into a single blended figure.
  • The date— every claim on this site is dated for this reason.
  • The country— a Toronto business gets different answers than the same query run from the US.
  • Named competitors— measured on the same question at the same time, so the comparison is like-for-like.

You can see this applied in the public engine snapshot 4 engines, one fixed prompt, captured June 30, 2026, each answer attributed to the engine that gave it.

03Where the numbers come from

Not all engines are equal as data sources.

This is the distinction most reporting in this category gets wrong. Only Microsoft Copilot reports citations back to publishers. Bing Webmaster Tools’ AI Performance report gives a first-party count: which of your pages were cited, how often, and the queries that retrieved them. ChatGPT, Perplexity, Gemini and Google AI Overviews expose no equivalent data to site owners.

So a citation countcan only ever come from Copilot. Everything else is observation — running the prompt and recording what came back. Both are useful; conflating them is not.

This is why the 14,600+ figure on this site is always attributed to Microsoft Copilot (Bing AI Performance) and never presented as a cross-engine total. The AI Citation Study publishes a sample of the underlying data (CC BY 4.0) covering April 25, 2026 to July 25, 2026, so the method can be checked rather than taken on trust.

04What this cannot prove

The limits, stated plainly.

  • Causation.Engines change their models and retrieval on their own schedule. A rise after an implementation is correlation with a plausible mechanism — not proof. A dated baseline is what makes it judgeable at all.
  • Completeness. Manual observation samples a fixed prompt set. It cannot cover every phrasing a real buyer might use.
  • Permanence. Any snapshot describes the day it was taken. Answers move.
  • Guaranteed outcomes. No one controls what an AI engine says. Anyone promising a guaranteed citation or ranking is describing something they cannot deliver.
05What stays private

Measurement is public. The fix is the product.

Everything above — how visibility is measured, recorded, and reported — is published so you can evaluate the work before paying for it. What is not published is the specific set of technical and structural changes that produced the results in the case study.

That is not evasiveness; it is the deliverable. You receive it in full, written down, in the AI Visibility Audit — documented so your own team can execute it, or so I can implement it for you.

06Questions

Why publish your methodology at all?

Because a number you can't interrogate isn't evidence. Most AI visibility providers report a score without saying which engines they queried, on what date, from which country, or how many times. That makes the result impossible to check and impossible to reproduce. Publishing the method is what separates a measurement from a marketing claim.

Which AI engines do you actually measure?

Six: ChatGPT, Perplexity, Google AI Overviews, Claude, Gemini, and Microsoft Copilot. They are not equivalent as data sources, which is the important part — only Copilot reports citations back to publishers, so it is the only engine where a first-party count exists. The rest are observed, not counted.

Do you use an AI visibility score out of 100?

No. A composite score hides more than it shows: two brands with the same score can have completely different problems. You get the underlying observations instead — where you appeared, where a competitor appeared instead, and what each engine said — because those are the things you can act on.

Can you prove your work caused a change in citations?

Not in the strict sense, and I won't claim otherwise. AI engines change their models and retrieval independently of anything a consultant does, so a rise after an implementation is correlation with a plausible mechanism — not proof of causation. What I do is record a dated baseline before any work starts, so at least the before-and-after is real and you can judge it yourself.

How do you avoid cherry-picking a good result?

By fixing the prompt set in advance, running it across engines rather than picking the flattering one, and recording the date and country of every answer. A single screenshot proves nothing — AI answers vary by session, phrasing, and region. Anything I report should say when it was captured and under what conditions.

See the method applied to your brand.

The free checker runs this same method on a small scale — your keywords, your competitors, dated and attributed.