An llm visibility audit template gives you a repeatable system for measuring exactly how often—and how accurately—AI models like ChatGPT, Perplexity, Claude, and Gemini surface your brand when buyers ask relevant questions. This case study documents a real audit we conducted for a B2B SaaS client in the project management space, including the exact prompt framework, scoring rubric, competitive gap findings, and the specific content interventions that produced measurable lift within 60 days.
The LLM Visibility Audit Template: Context and the Problem We Were Solving
The client—a 120-person project management SaaS company with $18M ARR—had noticed something troubling in their pipeline data. Demo requests from organic search were holding steady, but a new channel question they'd started asking in onboarding surveys revealed a growing cohort of users who said they first heard about the product through "an AI assistant." That cohort was growing at roughly 23% month-over-month through Q1 2026, yet the company had zero visibility into how AI models were representing their brand, their competitors, or their product category.
The stakes were significant. With roughly 34% of B2B software buyers reporting in early 2026 that they use AI assistants as a primary discovery channel before visiting vendor websites, an invisible or misrepresented brand in AI-generated answers is a top-of-funnel leak most companies aren't measuring. This client wasn't just invisible in spots—they were actively being described inaccurately in two of the four models we tested, with one model citing a pricing tier that was 18 months out of date.
Before we could fix anything, we needed a structured audit framework: a defined query set, a consistent scoring methodology, and a competitive baseline. Nothing the team had used for traditional SEO audits translated cleanly to this environment, because AI answer engines don't return rankings—they return narratives. That required an entirely different measurement approach.
"The most alarming finding wasn't that the brand was missing from AI answers—it was that when it did appear, 40% of those appearances contained factually incorrect information about pricing, integrations, or product capabilities."
What was explicitly NOT done at this stage: we did not attempt to reverse-engineer training data, we did not pursue any gray-hat prompt injection tactics, and we did not conflate traditional domain authority metrics with AI visibility scores. Those are different signals, and treating them as equivalent is one of the most common mistakes we see in this category of work.

Audit Strategy: The Prompt Framework and Scoring Rubric We Built
The core of any reliable LLM visibility audit is a prompt set that mirrors real buyer intent—not branded queries, but the category-level and problem-aware questions actual users type into AI assistants. We structured our prompt framework across three intent layers, with 12 prompts per layer for a total of 36 queries tested across four AI models (ChatGPT-4o, Perplexity, Claude 3.5 Sonnet, and Gemini 1.5 Pro), generating 144 response data points per audit cycle.
| Intent Layer | Example Query | What It Measures |
|---|---|---|
| Category Awareness (Layer 1) | "What are the best project management tools for remote teams?" | Brand mention rate in open-ended category responses |
| Problem-Aware (Layer 2) | "How do I manage dependencies across multiple projects without spreadsheets?" | Brand mention rate when buyer pain is foregrounded |
| Comparison/Decision (Layer 3) | "Compare [Brand] vs [Competitor A] for enterprise project tracking" | Accuracy of brand representation in direct comparisons |
Each response was scored against a five-dimension rubric: (1) Mention Presence (was the brand named at all?), (2) Mention Position (first, middle, or late in the response?), (3) Attribute Accuracy (were stated features, pricing, and integrations correct?), (4) Sentiment Framing (was the brand described neutrally, positively, or negatively relative to competitors?), and (5) Competitive Displacement (was a direct competitor recommended in place of the brand for a use case the brand clearly supports?).
Scores ran from 0–100 per dimension, producing a composite LLM Visibility Score (LVS) out of 500 per model, and an aggregate LVS out of 2,000 across all four models. For ongoing llm brand visibility tracking, this composite score gives you a single number you can trend over time without losing the dimensional detail that tells you what to fix.
The competitive baseline was established by running the same 36-prompt set and scoring rubric against three direct competitors. This revealed not just where the client was missing, but specifically which competitor was capturing those displaced mentions—critical for understanding where content investment would produce the fastest displacement reversal.
Implementation: Running the Audit Across Four AI Models
The audit ran over a structured two-week window in late February 2026. Each of the 36 prompts was executed in a fresh session (no conversation history) on each of the four models, with responses logged verbatim in a shared audit spreadsheet. All queries were run by two independent team members to control for session variability, with discrepancies reconciled through a third run where scores differed by more than 10 points on any dimension.
Tooling was deliberately minimal: no third-party AI monitoring platforms were used for the baseline audit, because we wanted a methodology that any team could replicate without a software budget. The core stack was Google Sheets (for logging and scoring), a shared prompt library in Notion (for consistency across testers), and Loom recordings of each session for dispute resolution. Total time investment for two auditors: approximately 28 hours across the two-week window.
Key findings from the baseline audit, quantified:
- Overall aggregate LVS at baseline: 612 out of 2,000 (30.6%)
- Brand mention rate across all 144 data points: 31% (the brand appeared in 45 of 144 responses)
- Of those 45 appearances, 40% contained at least one factual inaccuracy
- Competitive displacement rate (a competitor recommended instead of the brand for a supported use case): 58% of Layer 2 prompts
- Strongest model for brand mention: ChatGPT-4o at 44% mention rate; weakest: Gemini at 19%
With the baseline established, content interventions were mapped to specific scoring failures. Attribute Accuracy failures pointed to stale or thin technical documentation—the fix was refreshing the product feature pages and publishing a new integrations hub with structured, citable content. Competitive Displacement failures pointed to missing use-case content—the client had no published content directly addressing the 11 specific use cases where competitors were being recommended instead.
The intervention plan was sequenced across 60 days: weeks 1–2 focused on accuracy fixes (updated pricing page, refreshed feature documentation, new FAQ content addressing the specific inaccuracies found in audit responses); weeks 3–6 focused on use-case content (11 new long-form pages targeting the displaced use cases); weeks 7–8 focused on authority signals (updated third-party profiles, press release syndication to high-citation news sources, and structured data markup across key pages).
Results: Before-and-After Metrics After 60 Days of Content Interventions
The re-audit was run using the identical prompt set, scoring rubric, and two-auditor methodology exactly 60 days after the initial baseline. Results were meaningfully positive across every dimension, though the distribution of improvement was not what the team had predicted going in.
| Metric | Baseline (Feb 2026) | Re-Audit (Apr 2026) | Change |
|---|---|---|---|
| Aggregate LVS (out of 2,000) | 612 | 891 | +46% improvement |
| Overall brand mention rate | 31% (45/144) | 48% (69/144) | +17 percentage points |
| Factual accuracy rate (of appearances) | 60% accurate | 84% accurate | +24 percentage points |
| Competitive displacement rate (Layer 2) | 58% | 33% | -25 percentage points |
| Gemini mention rate (weakest model) | 19% | 36% | +89% relative improvement |
The most surprising finding in the re-audit: the accuracy interventions—updating the pricing page and feature documentation—drove proportionally more LVS improvement than the use-case content, despite the use-case content representing the larger production investment. Attribute Accuracy scores improved by an average of 38 points per model after the documentation refresh, while the 11 new use-case pages drove an average of 21 points per model on Competitive Displacement scores.
"Fixing what AI models were saying wrong about the brand turned out to be twice as high-leverage as creating new content to capture missing mentions. Accuracy is the foundation—everything else builds on it."
The pipeline impact was visible within the 60-day window, though the team appropriately treated this as directional rather than conclusive given the short timeframe: the "discovered via AI assistant" cohort in onboarding surveys grew from 23% month-over-month to 31% month-over-month in April 2026, and average deal velocity for that cohort was 11 days faster than the organic search cohort—suggesting that buyers who encounter accurate brand information in AI responses arrive at demos with a higher level of prior qualification.
What failed or underperformed: structured data markup (weeks 7–8 of the intervention plan) showed no measurable LVS impact at the 60-day mark. This doesn't mean it has no long-term value, but teams should not expect structured data changes to move AI visibility scores on short timelines. Similarly, third-party profile updates showed inconsistent impact—high on Perplexity (which actively cites sources), negligible on Claude and Gemini in this audit cycle.
Frequently Asked Questions
How many prompts should I include in an LLM visibility audit?
A minimum viable audit requires at least 18–24 prompts spread across the three intent layers (category awareness, problem-aware, and comparison/decision), tested across at least two AI models. The 36-prompt, four-model framework described in this case study (144 total data points) provides a statistically more reliable baseline, but the most important variable is consistency—use the same prompts every time you re-audit so you're tracking real change rather than query variation. Run each prompt in a fresh session with no prior conversation history to control for context contamination.
How long does it take to see results after making content changes for LLM visibility?
Based on this audit and similar engagements, accuracy-focused content changes (updating feature pages, pricing, and documentation) tend to show measurable LVS improvement within 30–60 days. New use-case content typically requires 60–90 days to show consistent lift, likely because AI models need time to encounter and integrate updated web content into their retrieval patterns. Structural changes like schema markup and third-party profile updates should be treated as longer-horizon investments with a 90–180 day measurement window.
What's the difference between an LLM visibility audit and traditional SEO auditing?
Traditional SEO audits measure ranking positions, crawl health, backlink profiles, and on-page optimization signals against search engine algorithms that return lists of URLs. An LLM visibility audit measures how accurately and frequently AI models generate brand mentions within natural-language responses—there are no rank positions, only narrative presence. The scoring dimensions (mention rate, attribute accuracy, sentiment framing, competitive displacement) have no direct SEO equivalents, which is why standard SEO audit templates don't translate to this use case.
Do I need specialized software to run an LLM visibility audit?
No—the baseline audit methodology described here was executed entirely with Google Sheets, a shared Notion prompt library, and direct access to the four AI model interfaces (ChatGPT, Perplexity, Claude, Gemini). Dedicated AI monitoring platforms can automate query execution and response logging at scale, which becomes valuable once you're running audits monthly across a large prompt set, but they are not required to produce a credible baseline. The discipline of consistent prompting and rigorous manual scoring matters more than tooling at the outset.
