Playbooks · Playbook
GEO Audit Playbook
A 1–2 day GEO audit playbook: prepare scope, capture AI answers, tag citations, prioritize gaps, and remeasure.
Run this playbook when leadership asks whether the brand is visible in AI answers but the team does not yet have repeatable evidence.
It is the operator version of the GEO audit workflow. The guide explains the sequence. This playbook tells a small team how to finish a first audit in one to two working days without turning the work into an open-ended research project.
Keep the scope narrow. One category, one market, one buyer role, and a short competitor list will produce a usable decision. A 40-prompt, five-market sweep will not, not on this timeline.
What this audit is for
A GEO audit answers four questions with evidence, not opinions:
- Which buyer questions currently produce answers about this category?
- Does the brand appear, and how is it framed?
- Which sources are cited, and do owned pages appear among them?
- Which competitors occupy the recommendation, comparison, or “safe default” slots?
Use Prompt, Presence, Proof, Position as the tagging language. Do not wait for a perfect measurement stack. A shared sheet plus stable prompts is enough for the first pass.
Day 0 to morning of day 1: prep
Do not start prompting until the inventory is written down. Prep usually takes 60–90 minutes.
| Prep item | Done when |
|---|---|
| Category and buyer | You can name one category, one market, and one primary buyer task. |
| Prompt themes | You have at least definition, category, comparison, recommendation, and problem-solution themes. |
| Competitor shortlist | You listed 3–5 names the buyer would actually compare, not every adjacent brand. |
| Owned sources | You listed the URLs that should be citable: category page, comparison page, docs, glossary, proof page. |
| Surfaces | You named the AI answer surfaces the team actually cares about this week. |
| Capture sheet | Columns exist for prompt, surface, date, answer summary, brand, competitors, citations, accuracy, action. |
If the owned-source list is empty, stop and collect it. An audit that cannot map an answer back to a page cannot produce a fix list.
Prompt set
Build 12–20 prompts for the first pass. Fewer than 10 leaves too many holes. More than 25 usually delays tagging.
Cover these jobs:
| Prompt job | Example shape | Why it belongs |
|---|---|---|
| Definition | “What is [category]?” | Shows whether the category is explained in language the team would accept. |
| Category discovery | “Best tools for [job]” | Shows who occupies the shortlist. |
| Comparison | “[Brand] vs [competitor]” | Shows framing, not only presence. |
| Recommendation | “Which [category] should a [role] use?” | Shows whether the brand is treated as a default, alternative, or omission. |
| Problem-solution | “How do I [buyer task]?” | Shows whether owned how-to sources can enter the answer. |
| Accuracy trap | “Does [brand] support [claim]?” | Shows stale or invented product descriptions. |
Write the exact wording in the sheet and freeze it for this audit. Small wording changes are a later experiment, not part of the first capture. See prompt tracking for why the string itself is evidence.
Capture
Run the same prompt set across the named surfaces. Capture on the same day if possible so the batch is comparable.
For every row, save:
- exact prompt wording;
- surface or product name;
- capture date and time;
- a short answer summary plus the claims that matter;
- visible citations or source names;
- whether the brand is mentioned, omitted, or misdescribed;
- which competitors appear, and in what order if a list is given.
Do not score yet. Capture first. Scoring during capture usually collapses the record into a vibe.
If a surface shows no citations, record that as a finding. Missing attribution is still evidence; it just changes how much weight the row can carry for source work.
Tag mentions, citations, and competitors
Tagging is the actual analysis. Budget most of day 1 afternoon for it.
Use a consistent code, then write one sentence of context:
| Tag | Meaning |
|---|---|
| Present / omitted / misdescribed | Brand presence in the answer body. |
| Cited owned / cited third-party / no citation | Whether a source is visible, and whose it is. |
| Lead / co-mentioned / alternative / absent | Competitor or brand position in the recommendation frame. |
| Accurate / stale / invented | Whether the answer matches current public facts the team can defend. |
Then connect each weak row to a page or proof gap:
- omitted on category prompts often means the category page does not state the job clearly;
- present but uncited often means the page is vague, undated, or hard to extract;
- cited competitor docs often means owned comparison or specification pages are thin;
- misdescribed claims often means public pages, partner copy, or third-party listings disagree.
Use the content citability checklist when the issue is extractability rather than missing coverage.
Prioritize fixes
Do not produce a 30-item content backlog. Rank a short action list that a team can start this week.
Score each issue on:
- Prompt importance: is this a buyer question the team already pays to win in search?
- Frequency: does the gap repeat across prompts or surfaces?
- Fixability: is there an owned page, proof asset, or entity-consistency issue the team can change?
- Risk: is the answer actively wrong about product, pricing, or category?
A practical priority order:
| Priority | Typical pattern | First fix |
|---|---|---|
| P1 | High-intent recommendation or comparison prompts omit or misdescribe the brand | Repair the owned category, comparison, or product-fact page |
| P2 | Brand appears, but citations go to competitor or generic sources | Strengthen extractable claims, dates, and source clarity on owned URLs |
| P3 | Definition prompts describe the category loosely | Publish or tighten a durable explainer that matches buyer language |
| P4 | One-off odd answers with no repeat pattern | Log for the next measurement pass; do not rebuild the site around them |
The output of this step is the expected measurement outcome: a ranked list of prompts, citation gaps, competitor visibility patterns, and source improvement tasks.
Re-measure
Re-measure only after a defined change, not after “we talked about GEO.”
Wait until at least one P1 or P2 page has been updated, then rerun the same prompt set on the same surfaces. Keep the original rows. Add a second capture date. Compare presence, citations, framing, and accuracy notes.
If nothing changed, that is still a result. It usually means the wrong page was edited, the claim is not corroborated off-site, or the prompt set was too brand-centric. Fold that lesson into how to measure AI visibility rather than widening the audit immediately.
Output format
Create one audit table with these columns:
| Column | Use |
|---|---|
| Prompt | Frozen wording from the set |
| Surface | Named AI answer surface |
| Answer summary | The claim that matters, not the full transcript |
| Brand presence | Present, omitted, or misdescribed |
| Competitor presence | Who else appears, and in what frame |
| Citations | Owned, third-party, or none |
| Accuracy notes | Stale, invented, or acceptable |
| Recommended action | One page, proof, or entity fix, with a priority |
Attach a one-page memo: scope, date range, top five findings, and the remeasure date. That package is enough for an SEO operator, a growth lead, or an agency client review.
Operator checklist
- Scope is one category, not the whole site.
- Prompt set covers definition, category, comparison, recommendation, and problem-solution jobs.
- Every row has a date, surface, and frozen prompt.
- Mentions, citations, and competitor frames are tagged separately.
- Actions point to URLs, not slogans.
- Remeasurement uses the same prompts after a real source change.
If the team later needs a longer validation cycle rather than a first snapshot, move to the 90-day operating rhythm. If the question is which measurement jobs tools should support, use the 2026 tool landscape.