Guides · Guide
GEO Audit Workflow
An operator workflow for scoping, capturing, tagging, and acting on GEO audits of brand visibility, citations, and answer quality.
A GEO audit checks how AI answer systems describe a brand, category, and competitors today. The useful output is evidence: what the answers say, which sources they cite, where the brand is missing or misrepresented, and which pages can change that.
Use the GEO audit playbook for a first-run table. Use how to measure AI visibility when the audit should become a recurring program.
1. Scope the audit
Do not start with every prompt and every surface. Record these decisions before capture:
| Scope item | Decision to record |
|---|---|
| Category | The market or problem space the buyer is researching. |
| Brand and product names | Exact names, former names, and names to ignore. |
| Competitor set | A shortlist, not every adjacent vendor. |
| Markets and languages | Country, language, and whether localized answers are in scope. |
| Surfaces | Which AI search, answer engine, or assistant experiences will be captured. |
| Timebox | How many prompts, how many surfaces, and when capture closes. |
A focused first audit often covers one category, 20 to 40 prompts, two or three surfaces, and five to eight competitors. Broader programs come after the first evidence set is usable.
Counterexample: collecting twenty random chats from different people, then reporting “we are not visible.” That is a conversation. The prompt, surface, date, and competitor set were never controlled.
2. Build the prompt set
The prompt set is the measurement object. Prompt tracking only works if the questions reflect real discovery and evaluation work.
Cover these prompt types:
| Type | What it tests | Example shape |
|---|---|---|
| Definition | Category language and whether owned sources explain the topic. | ”What is generative engine optimization?” |
| Category | How the market is framed and which tool types appear. | ”What are GEO tools?” |
| Recommendation | Brand presence, order, and shortlist quality. | ”Which AI visibility tools should a B2B SaaS team consider?” |
| Comparison | How alternatives are described together. | ”GEO tools vs SEO rank trackers” |
| Workflow | Whether owned guides appear as practical sources. | ”How should an agency run a GEO audit?” |
| Problem-solution | Whether the brand is attached to the job, not only the category name. | ”How do I check if an AI answer cites my site?” |
Keep core prompts stable. Put new buyer phrasing in a backlog so the trend is not destroyed by constant rewording. SEO demand is an input, not a keyword paste: map impression-heavy queries into the same intent families, then add conversational variants.
3. Capture protocol
Capture is a protocol, not a screenshot dump:
- Use the recorded prompt wording. Do not paraphrase mid-run.
- One prompt per record. Do not merge follow-up chat unless that follow-up is a tracked prompt.
- Record surface, logged-in or logged-out state, market, and language.
- Save the full answer text and visible citations exactly as shown.
- Record the capture date. Do not edit the answer before tagging.
If a surface shows no citations, still keep the answer. Uncited answers are evidence. They often reveal framing problems that citation counts miss. See citation tracking.
4. Tagging taxonomy
Tagging turns answer text into comparable fields. Use a small, consistent taxonomy.
| Tag | Values | How to use it |
|---|---|---|
| Brand presence | mentioned / omitted / unclear | Unclear is for ambiguous nicknames or product-line confusion. |
| Mention role | recommended / compared / defined / passing / negative | A passing mention is not a win. |
| Position | first / middle / last / not applicable | Use only when the answer ranks options. |
| Accuracy | accurate / stale / incomplete / incorrect / mixed | Mixed is common; do not force a single score. |
| Citation owner | owned / competitor / third-party / uncited | Track the supporting source, not only the brand name. |
| Competitor hit | which shortlist names appear | Relative visibility is the audit, not an isolated mention. |
Do not invent extra tags during the first audit. If a new distinction keeps appearing, add it next cycle and re-tag the baseline.
Example: the brand is third in a five-name shortlist, described with last year’s category, and cited only via a partner blog. Presence is true. The action is to fix category language and owned-source citability, not to celebrate the mention.
5. Evidence fields
Keep enough fields that a later run can be compared without tribal knowledge.
| Field | Why it is required |
|---|---|
| Prompt ID and wording | Ties the record to the controlled set. |
| Intent family | Groups definition, category, comparison, and workflow prompts. |
| Surface, market, language, date | Makes runs comparable across systems and locales. |
| Full answer text | Scores cannot reconstruct claims. |
| Brand mention and role | Separates presence from recommendation quality. |
| Competitor names and order | Explains relative visibility. |
| Cited URLs and source owner | Shows who is supporting the answer. |
| Accuracy notes | Records stale, missing, or unfair claims. |
| Follow-up action | Connects evidence to a page, source, or entity fix. |
If a field cannot be filled, write unknown rather than guessing. Unknown capture conditions are themselves an audit finding.
6. Output an action list
The audit is finished when it produces an action list, not when it produces a score.
Group actions by the thing an editor or SEO can change:
| Finding | Typical action |
|---|---|
| Brand omitted on category and recommendation prompts | Clarify category pages, comparison assets, and entity language. |
| Brand mentioned, owned pages never cited | Improve citability: definitions, evidence, headings, and source freshness. |
| Third-party pages define the category | Strengthen owned glossary and guide pages that answer the same question. |
| Competitors recommended first with better comparison pages | Build or refresh a comparison that names alternatives fairly. |
| Answer is accurate but stale | Update product, category, and positioning language across owned sources. |
| Answer is incorrect | Fix the owned source of truth, then check which third-party pages still repeat the error. |
Each action should name an owner, a URL or source type, and the prompt family it is meant to change. “Create more content” is not an action.
7. Cadence
A first audit is a baseline. Re-run the same core prompt set after the first action batch lands.
| Situation | Practical cadence |
|---|---|
| First baseline | One complete capture window, then a re-run after the first fixes. |
| Stable category, small team | Monthly review of the core set. |
| Competitive category or frequent product changes | Weekly or biweekly on recommendation and comparison prompts. |
| After a launch, repositioning, or major content refresh | Immediate targeted re-run, then return to the regular cadence. |
Changing the prompt set every cycle destroys the trend. Add prompts in a named backlog and promote them into the core set only when they will be kept.
Common audit failures
- Treating one chat answer as the market.
- Mixing logged-in, personalized, and logged-out captures without recording the condition.
- Counting mentions without tagging role, position, or accuracy.
- Ignoring uncited answers because they do not produce a URL.
- Auditing only brand-name prompts, which overstates visibility.
- Writing a score with no action list.
- Changing prompts, surfaces, and competitors between the baseline and the re-run.
- Turning the report into a tool pitch instead of source and content work.
Next step
Re-run from the GEO audit playbook table, then keep the loop in how to measure AI visibility and citation tracking.