Frameworks · Framework
AI Visibility Measurement Framework
Define Prompt, Presence, Proof, and Position so AI visibility work produces evidence, not anecdotes.
Use Prompt, Presence, Proof, Position to keep SEO-informed AI visibility measurement practical. The four parts are an operating loop, not a slogan: define the buyer questions, check whether the brand appears, record the citations and claims, then compare how the answer frames competitors.
This framework keeps GEO work tied to observable answer evidence. It is the measurement language behind how to measure AI visibility and the GEO audit playbook.
The four parts at a glance
| Part | Operating question | Required evidence | Primary metrics | Typical failure mode |
|---|---|---|---|---|
| Prompt | Which buyer questions are we measuring? | Frozen prompt set, intent labels, surfaces, cadence | Prompt coverage, prompt stability | Measuring brand-name vanity queries only |
| Presence | Does the brand appear in the answer? | Answer text, mention tag, date, surface | Mention rate, AI share of voice | Treating one chat as a benchmark |
| Proof | What sources and claims support the answer? | Visible citations, owned vs third-party URLs, accuracy notes | Citation rate, citation quality | Counting links without checking the claim |
| Position | How is the brand framed against alternatives? | Competitor co-mentions, list order, role in the answer | Recommendation position, alternative rate | Reporting “we were mentioned” while the answer recommends someone else |
Prompt
Prompt is the unit of measurement. Pages still matter, because they are often the sources behind an answer, but the buyer question is what the system is answering.
A prompt set should include definition, category discovery, comparison, recommendation, and workflow questions. Brand-only prompts (“What is [brand]?”) can sit in a small control group. They should not be the whole set. If the team only asks questions that already contain the brand name, Presence will look healthier than category reality.
Required evidence:
- exact wording;
- topic cluster and buyer intent;
- language and market if those change the job;
- the surfaces that will be sampled;
- the date the wording was frozen.
Failure modes:
- rewriting prompts every week and calling the movement a trend;
- mixing several buyer jobs into one oversized question;
- skipping comparison and recommendation prompts because they feel “too competitive.”
If the prompt set is unstable, none of the later metrics are interpretable. Prompt tracking is the discipline that protects the rest of the framework.
Presence
Presence asks whether the brand is in the generated answer at all, and whether the mention is accurate.
Required evidence:
- the answer text, or a faithful excerpt of the claims that mention the brand;
- a tag: present, omitted, or misdescribed;
- surface and capture date;
- reviewer notes when the mention is partial (“the product category is right, the packaging claim is wrong”).
Useful metrics:
- mention rate across the prompt set;
- mention rate by prompt job (definition vs recommendation);
- AI share of voice against the competitor shortlist;
- volatility: how often presence flips between capture windows.
Failure modes:
- scoring Presence from memory after a demo;
- collapsing “mentioned once in a list of twelve” into the same win as a featured recommendation;
- ignoring misdescription because the brand string appeared.
Presence without Proof and Position is a vanity signal. A brand can be present and still be uncited, outdated, or framed as a weak alternative.
Proof
Proof is the evidence layer: citations, source attribution, and claim quality.
Required evidence:
- visible citations or named sources, including an explicit “none shown” record;
- owned URL vs third-party URL vs competitor URL;
- whether the cited page actually supports the claim;
- accuracy notes for product facts, pricing, geography, and category membership.
Useful metrics:
- citation rate for owned sources;
- citation quality: relevant, current, and claim-aligned, not merely linked;
- unsupported-claim rate: answers that make specific statements with no visible source;
- stale-description rate: answers that contradict the current public page.
Failure modes:
- treating any citation as a win, including citations to thin listicles or outdated directories;
- optimizing only for being cited while the cited page is the wrong one;
- skipping accuracy review because the citation count looks busy.
Proof is where answer monitoring meets source work. If owned pages never appear, the next action is usually extractability, entity consistency, or missing corroboration, not a new dashboard.
Position
Position is competitive framing. It asks where the brand sits in the answer’s recommendation structure.
Required evidence:
- which other brands, products, or “generic category” options appear;
- order or grouping when the answer produces a list;
- the role assigned to the brand: default pick, specialist, alternative, warning, or omission;
- the comparison criteria the answer used, even if those criteria are crude.
Useful metrics:
- recommendation position (first, later, alternative-only, absent);
- co-mention rate with named competitors;
- win/loss on comparison prompts;
- criterion mismatch rate: the answer ranks on a job the brand does not actually compete for.
Failure modes:
- celebrating Presence on a comparison prompt where the brand is the example of what not to buy;
- averaging Position across unrelated prompt jobs;
- changing competitor names between waves so share of voice cannot be compared.
Position is the metric leadership usually thinks it is buying when it asks “are we visible?” Make the frame explicit.
How the four parts work together
Run them in order, then loop:
- Freeze Prompt.
- Capture Presence.
- Attach Proof to each row.
- Score Position against the same competitor shortlist.
- Turn gaps into page and source tasks.
- Remeasure the same prompts after a defined change.
A row is only complete when all four parts have evidence. Mention rate without citations cannot explain why. Citations without Position cannot explain whether the answer helps. Position without a frozen prompt cannot be trended.
What “good” looks like in the first 30 days
A small program is working when the team can show:
- a prompt set of 12–20 questions covering more than brand lookups;
- a capture log with dates and surfaces;
- mention, citation, and position tags that a second reviewer could repeat;
- a short fix list mapped to URLs;
- a planned remeasure window after those URLs change.
It is not working when the only artifact is a single composite score, a screenshot from one chat, or a slide that says visibility “feels better.”
Where to use this next
Use the GEO audit playbook to run the first 1–2 day snapshot with this language. Use how to measure AI visibility to keep the cadence after the snapshot. If the team is choosing measurement jobs rather than writing copy, read the 2026 AI visibility tool landscape.