AI visibility metrics belong on a report in the shape of the new customer journey, and almost nobody does that. Discovery, shortlist, visit, evaluation, intent, purchase. The first three stages now happen inside someone else's interface, so opening with referral sessions makes the channel look like a rounding error at 1.08% of web traffic. Opening at discovery, with citation share against competitors, tells a usable story from the same data.
Map AI marketing KPIs to the new customer journey
The buyer no longer starts on a search results page. They ask an assistant. They read a shortlist. They form a view before they ever reach your site. Six metric categories track that journey, one per stage. A report that opens with referral sessions has skipped the first two.
The old funnel let you watch every step. The new one does not. The first two stages happen inside someone else's app, which is why visible AI clicks are about 1.08% of web traffic while the influence behind them is much larger.
| What changed | Search era | AI answer era |
|---|---|---|
| What the buyer types | Keywords. Two or three words, stripped of context | A prompt. A full question with their situation in it |
| What they get back | Ten links to choose from | One answer with a short shortlist |
| Where they compare | Across several open tabs, after the click | Inside the answer, before any click |
| How many options | As many as they open tabs for | Three to five, and the first is picked most |
| What you can see | The query, the rank, the click | No query, no rank, and often no click |
| Where they evaluate | Across your site, review sites and competitor tabs | Inside the answer, before any site is opened |
| What you optimise | Ranking for a keyword | Being cited in an answer |
The evaluation moved too, and that is the part most teams miss. In the search era a buyer opened your page, a competitor's page and a review site, then compared them across tabs. Your pricing page did the work. Your comparison page did the work. Now the assistant does that work. It reads all three sources, weighs them, and hands back one shortlist with the comparison already made.
So the pages you built to persuade are read by a model before a human sees them, and often instead of a human seeing them. That is why being cited in an answer is now the thing you optimise for. You cannot rank in an answer. You are either in it or you are not, so the first metric is how often you appear.
| Journey stage | What the buyer does | The metric category that sees it |
|---|---|---|
| Discovery | Asks an assistant, reads a shortlist | Visibility. Citation share, brand mention rate, competitor co-mention |
| Shortlist | Forms a view from how you are described | Quality. Sentiment, factual accuracy, crawler reachability |
| Visit | Clicks through, or searches your brand later | Traffic. AI referral sessions, branded search, landing page mix |
| Evaluation | Explores the product and the pricing page | Engagement. Bounce, duration, pages per session |
| Intent | Signs up, requests a demo, enters pipeline | Pipeline. AI-sourced opportunities, win rate, cycle length |
| Purchase | Converts to a paid plan | Revenue. Paid conversions, new ARR, CAC payback |
Share of voice in AI search replaces the old rank report
Position fails as a stable measurement, so any dashboard reporting it is selling noise. SparkToro and Gumshoe.ai ran 2,961 prompts with 600 volunteers and found the odds of the same brand list appearing twice were under 1 in 100. The same list in the same order came in near 1 in 1,000.
| Metric | How to calculate | What it tells you |
|---|---|---|
| Citation share | Prompts where you are cited ÷ prompts tracked | Presence with a linkable source |
| Brand mention rate | Answers naming you ÷ answers generated | Influence beyond citations. ChatGPT mentions run about 3.2× citations |
| Prompt coverage | Prompts where you appear ÷ mapped buyer prompts | Breadth across the buying journey |
| Average citation position | Mean rank within the answer shortlist | Directional only. Never report as a rank |
| Competitor co-mention | Your answers also naming rival X ÷ your answers | The comparison set engines placed you in |
| Model variance | Citation share by engine, plus cross-engine overlap | Only about 2.4% of cited URLs overlap across engines |
AI referral traffic misses most of your visits
Every number here is a floor, so say so on the slide. The real figure is higher and you cannot see it. Similarweb found 55.9% of AI-influenced visits arrive as branded search rather than a trackable referral, so the visit stage structurally under-reports.
| Metric | Source and calculation | Benchmark |
|---|---|---|
| AI referral sessions | GA4 sessions matching the AI channel regex | The measurable click floor |
| AI referral share | AI sessions ÷ total sessions | 1.08% average across 13,770 domains. IT 2.80%, Utilities 0.35% |
| Growth rate | Month over month change in AI sessions | About 1% month over month on the Conductor panel |
| Landing page mix | AI sessions grouped by landing page | Product and comparison pages usually outrank blog |
| Engine mix | AI sessions by referrer host | ChatGPT is 87.4% of AI referrals |
Never report crawler hits as traffic. Cloudflare's crawl-to-refer ratios ran as high as roughly 70,900 crawls per referred visitor, so mixing bots into this stage inflates it by orders of magnitude.
Be honest about the AI search conversion rate evidence
Published multiples range from no significant difference to 23×, and the most disciplined study sits at the bottom. Amsive analysed 54 sites over six months: LLM referrals converted at 4.87% against organic at 4.60%, with a paired t-test returning p = 0.794. The B2B subset ran 2.03% against 1.68%, also not significant.
| Metric | Calculation | What is known |
|---|---|---|
| Bounce, duration, pages | Standard engagement, segmented to the AI channel | Adobe retail data: 48% longer sessions, 13% more pages, 33% lower bounce |
| Conversion rate vs organic | AI channel CVR ÷ organic CVR | Range from parity to 23×, depending on sample and conversion event |
| AI-sourced signups | Signups carrying an AI origin flag | Ahrefs: 0.5% of traffic drove 12.1% of signups |
| Lead to SQL rate | SQLs ÷ AI-sourced leads | Varies 20% to 54% by vertical on published B2B panels |
Plan on 4× to 10× for B2B, hold it loosely, and say where each figure came from. The range comes from differences in sample size, vertical and what counts as a conversion: a software trial start is not a retail purchase.
Pipeline attribution is what wins you more marketing budget
Every earlier stage points at this one. Report AI-sourced and AI-influenced pipeline created, win rate, average deal size and cycle length, each split by first-touch source.
| Metric | Calculation | Why it matters |
|---|---|---|
| AI-sourced pipeline | Sum of opportunity value where the AI origin flag is set | The headline number for a CRO |
| AI-influenced pipeline | Adds opportunities where self-reported source names an AI tool | Recovers the no-click majority |
| Win rate by source | Closed-won ÷ opportunities, split by first touch | Reveals quality that lead counts hide |
| Average deal size | Closed-won value ÷ deals, split by first touch | Fewer, larger deals can beat higher volume |
| Sales cycle length | Median days from first touch to close | Pre-qualified buyers should close faster |
| CAC by source | Channel cost ÷ customers acquired | Makes the ROI question answerable |
The mechanics of joining a source to a closed deal are in revenue attribution.
Revenue attribution metrics count paying customers
Define revenue before you measure it. Here it means a new customer on a paid plan. That is the moment a card is charged for the first time. A trial start does not count. Nor does an MQL, or pipeline created.
The definition matters because AI search sits at the start of the journey. A signup proves the citation was seen. A first payment proves it was worth something. Every metric below is anchored to the payment.
| Metric | Calculation | Caveat |
|---|---|---|
| AI-sourced paid conversions | New customers on a paid plan carrying an AI origin flag | The primary number. Count customers rather than signups |
| Signup to paid rate | AI-sourced paid conversions ÷ AI-sourced signups | Compare against your blended rate rather than an industry figure |
| Time to first payment | Median days from AI-sourced signup to first charge | Pre-qualified buyers should convert faster |
| New ARR from AI-sourced customers | Sum of first-year contract value on those accounts | State the attribution model behind it |
| Revenue per AI session | AI-sourced new ARR ÷ AI referral sessions | Denominator is a floor, so the figure runs high |
| Revenue per citation | AI-sourced new ARR ÷ citations in period | Requires a closed loop. Rare in practice |
| CAC payback on AI-sourced | Programme cost ÷ new ARR from AI-sourced customers | Published 3× to 8× ROI figures are vendor-sourced |
| Incremental lift | Paid conversions against a matched holdout | The only causal measure, and rarely feasible |
Both motions need attribution, and they fail in different ways. In a product-led, self-serve motion the whole chain is observable: AI answer, signup, activation, first charge. It can close in days, so the origin flag rarely has time to get lost.
Sales-led is harder. The same buyer takes months, several people from the account get involved, and the person who signs the order is almost never the person who asked the assistant. Roll touches up to the account level before assigning credit, otherwise the channel that created the demand looks like it contributed nothing.
At current volumes most companies cannot produce a statistically significant per-citation revenue figure. Say that in the first meeting rather than being forced into it later. What you can defend is the AI-sourced customer count and the directional trend.
AI visibility tracking tells you what is coming next
Branded search predicts revenue better than clicks do. Similarweb found AI-recommended brands were 2.5× more likely to get a site visit within seven days, with 55.9% arriving as branded search. Ahrefs' correlation work put branded web mentions against AI visibility at 0.50 to 0.74, while backlinks and ad spend sat below 0.30.
Track citation share weekly, then test correlation with lag against branded search volume in Search Console, direct traffic in GA4, and net-new pipeline. When citation share rises and branded search follows within a sales cycle, you have a defensible story. When it rises and nothing follows, the citations are low quality, and that diagnosis is worth more than a precise number.
Attribution reporting the board will believe
One slide. Six stages. Show movement rather than scores. Tie every metric back to pipeline. State the method behind each figure and show the gap between self-reported and tracked attribution rather than resolving it.
- Citation share and mention rate against named competitors, with confidence intervals.
- Mention sentiment and accuracy, because being named badly is worse than not being named.
- AI referral sessions, labelled as a floor every time.
- Branded search and direct trend, which is where the invisible influence lands.
- AI-sourced and AI-influenced pipeline, with the reconciliation gap named.
The threshold worth watching: when AI-sourced and AI-influenced pipeline passes roughly 5% of new pipeline, or branded search rises in step with citation share, the programme has earned a larger allocation. Below that, hold at maintenance spend and say so.
What a CMO dashboard for AI search looks like
The clearest public template for this is Kevin Indig's AI visibility ladder, published in Kyle Poyar's Growth Unhinged in July 2026. Indig uses it with Airbnb, Asana and Xero, and it borrows its structure from Andy Grove's paired leading and lagging indicators in High Output Management. The logic is simple: when revenue attribution lags, you want to know whether you are on track long before the revenue arrives.
His ladder runs leading indicators, then quality guardrails, then lagging indicators, mirroring a Retrieved, Cited, Trusted progression. Three setup rules make the dashboard trustworthy:
- Freeze 20 to 50 high-intent prompts across personas, use cases and buying stages, for at least four weeks, so you measure real change rather than prompt drift.
- Log every run: prompt, model, location, answer, cited URLs, brands mentioned and shortlist position. That table is the raw material every rung reads from.
- Run two clocks. Weekly the team checks signal quality. Monthly the CMO checks allocation.
Indig names three traps worth quoting directly: vanity metrics treated as the destination, false precision from decimals on a number that re-rolls monthly, and mixing leading indicators with outcomes without a model connecting them. His framing of last-click AEO measurement is the sharpest in the literature: valuing it by referral clicks is like valuing a Super Bowl ad by QR-code scans.
Every row needs a decision and an instrument
A number that triggers no decision, or that you have no way to move, is the first trap on that list. The scorecard above is only worth presenting if each row, on every rung, comes with two more columns: what you do when it moves, and what has to be instrumented before you can move it. Without both, the dashboard is a weather report.
| Metric | The decision it triggers | The instrument required |
|---|---|---|
| Leading indicators | ||
| Bot crawls (Search crawlers only) | Unblock or whitelist search crawlers if drop exceeds 20% | Cloudflare Radar / server access logs filtered to AI user agents |
| Citation share on frozen prompt set | Investigate which third-party reviews and comparison sites took citations | Automated prompt sampling across 10+ engines, frozen prompt corpus |
| Share of voice vs named rivals | Adjust content allocation toward topics where competitors dominate citations | Competitor entity co-mention tracking in same answers |
| Quality guardrails | ||
| Sentiment of the mention | Flag negative sentiment to product and PR teams for messaging correction | NLP sentiment classification on answer extracts |
| Factual accuracy | Issue schema updates and publish correction content for hallucinated claims | Entity-attribute extraction comparing answers against pricing/feature specs |
| Crawler reachability | Fix robots.txt, paywalls or dynamic rendering blocking AI search bots | Daily HTTP 200 checks on robots.txt and sitemap URLs for GPTBot, PerplexityBot |
| Lagging indicators | ||
| Branded search volume | Attribute top-of-funnel lift to AI discovery when correlated with citation share | Google Search Console API branded query tracking with lag analysis |
| AI referral sessions | Analyze landing page conversion paths for high-intent referral cohorts | GA4 custom channel grouping for AI referrers + UTM taxonomy |
| AI-sourced signups | Compare AI signup rate vs blended organic to calibrate channel quality | First-party origin cookie captured in hidden form field at registration |
| AI-sourced paid conversions | Scale or trim AEO budget based on customer acquisition cost (CAC) | CRM contact-to-deal join with read-only first-touch origin field |
| New ARR from AI-sourced | Report ROI to the board; expand AEO budget when ARR exceeds 5% threshold | Stripe / billing integration joined to CRM closed-won deal IDs |
Read the table as a build list. Any row where the instrumentation column is not yet true is a row you should leave off the scorecard entirely. A dashboard with three populated, actionable rows outperforms one with twelve uninstrumented scores.
AI Visibility & Revenue Attribution Scorecard
Scope: 10 Generative AI Engines · 50 Frozen Prompt Corpus · HubSpot CRM & Stripe Billing Joined
| Metric & Scope | Cadence | Target | Current MTD | MoM Trend | Status | Action Triggered |
|---|---|---|---|---|---|---|
| Tier 1 · Leading Engine Signals (Early Visibility & Crawl Health) | ||||||
| Bot Crawl Volume AI search crawlers (GPTBot, Perplexity) | Weekly | > 10,000 | 14,280 | +18.4% | HEALTHY | Monitor CDN logs; maintain bot whitelisting |
| Citation Share Frozen 50-prompt buying corpus | Weekly | > 25.0% | 28.4% | +4.2% | ON TRACK | Expand category comparison & integration docs |
| Share of Voice vs Named Rivals Primary 3 SaaS competitors | Weekly | > 1.50x | 1.82x | +0.32x | OUTPERFORMING | Defend tier-1 feature prompts against competitor shifts |
| Tier 2 · Quality & Accuracy Guardrails (Damage Prevention & Sentiment) | ||||||
| Mention Sentiment Net Sentiment Score (-100 to +100) | Real-Time | > +65 | +74 | +6 pts | POSITIVE | Escalate negative snippets to product & PR |
| Factual Accuracy Rate Verified pricing & feature claims | Bi-Weekly | > 95.0% | 96.8% | +1.8% | VERIFIED | Publish structured schema updates for hallucinated pricing |
| Crawler Reachability HTTP 200 success on key endpoints | Daily | 100.0% | 99.9% | 0 errors | PROTECTED | Audit robots.txt and sitemaps daily |
| Tier 3 · Lagging Outcomes & Board Metrics (Pipeline & Closed ARR) | ||||||
| Branded Search Volume GSC impressions (AI discovery lift) | Monthly | +10.0% YoY | 48,500/mo | +14.3% | HIGH LIFT | Attribute top-of-funnel halo to generative engines |
| AI Referral Sessions GA4 custom AI channel grouping | Monthly | > 3,000 | 3,840 | +22.1% | SURPASSED | Optimize landing page conversion paths for AI visitors |
| AI-Sourced Signups First-touch origin cookie join | Monthly | > 140 | 162 signups | +15.7% | EXCEEDING | Calibrate CAC against blended organic channels |
| AI-Influenced Pipeline CRM opportunities with AI touchpoint | Monthly | > $1.20M | $1.42M ARR | +18.3% | ABOVE PLAN | Feed AI query intent data to enterprise AE teams |
| Recognized Net-New ARR Stripe closed-won AI revenue | Monthly | > 5.0% total | $285k (7.4%) | +1.9% share | BOARD VALIDATED | Reallocate budget: expand 2027 AEO investment |
What to ask AI visibility tools before you buy one
Most of the category measures presence in answers and stops at the citation. Three questions separate a dashboard from an attribution system.
- Can you show me a closed-won deal and the AI answer that started it? Almost nobody in the category can answer this.
- How many prompts, runs and engines does my tier cover? Most price on tracked prompt volume in tiers of 15, 100 or 500. Given the variance above, a small set run once is meaningless, so entry tiers sell an underpowered sample.
- Do you report ranking position? If yes, they have not read the research on non-determinism.
How ThriveStack citedby differs from other AEO tools at every stage
This metric set is only useful if the last stage is populated, and the last stage is paid customers.
citedby is the AI visibility product inside ThriveStack. It samples answers across roughly ten engines and tracks which of your pages get cited. It then carries that origin into your CRM and billing data, so a citation can be traced to a closed deal. The loop it runs is analyze, fix, attribute.
- Citation share, mention rate and co-mention on a frozen prompt set across roughly ten engines, reported as movement with confidence intervals.
- Sentiment and factual accuracy of each mention, so quality sits beside volume in one view.
- AI referral capture covering the engines GA4's native channel omits, including Perplexity and Claude.
- Branded search and direct correlated against citation share with lag, recovering the influence that lands elsewhere.
- AI-sourced pipeline and closed-won revenue through CRM and billing integration, which is the stage that renews budget.
Build this report from your own data
citedby tracks citation share, mention quality, AI referrals and AI-sourced pipeline in one view, connected to your CRM and billing data.
AI visibility metrics: FAQ
What AI visibility metrics should a CMO track?
Six categories, one per stage of the new customer journey. Discovery: citation share, mention rate, co-mention. Shortlist: sentiment, factual accuracy, crawler reachability. Visit: AI referral sessions and branded search. Evaluation: bounce, duration, conversion rate. Intent: AI-sourced opportunities, win rate, deal size. Purchase: paid conversions and new ARR. Report movement against named competitors rather than absolute scores.
Should AI ranking position be on the report?
No. SparkToro and Gumshoe.ai found under a 1 in 100 chance of an AI returning the same brand list twice across 2,961 runs, and roughly 1 in 1,000 for the same order. Position in AI answers moves too much to measure.
What is the difference between citation share and mention rate?
Citation share counts answers where you appear with a linkable source. Mention rate counts answers naming you at all, with or without a link. BrightEdge measured about 2.37 ChatGPT mentions per response against 0.73 citations, so tracking citations alone misses most of the influence.
When does AI visibility justify more budget?
When AI-sourced and AI-influenced pipeline exceeds roughly 5% of new pipeline, or when branded search rises in step with citation share over a full sales cycle. Below that, hold at maintenance spend rather than arguing from leading indicators alone.
Which AI visibility metric best predicts revenue?
Branded search lift. Similarweb found AI-recommended brands were 2.5 times more likely to receive a visit within seven days, with 55.9% arriving as branded search. Ahrefs put branded web mentions against AI visibility at 0.50 to 0.74 correlation, while backlinks and ad spend sat below 0.30.
What goes on a CMO dashboard for AI visibility?
Three rungs, following Kevin Indig's AI visibility ladder in Growth Unhinged, July 2026. Leading indicators: bot crawls, citation share, share of voice. Quality guardrails: sentiment, factual accuracy, crawler reachability. Lagging indicators: branded search, AI referrals, AI-sourced signups and paid conversions. Report movement across the rungs rather than a single AEO score.
What counts as revenue in AI search attribution?
A new customer converted to a paid plan, meaning the first charge rather than a trial start, an MQL or pipeline created. AI search operates at the discovery end of the journey, so a signup only proves the citation was seen. The first payment proves it was worth something.
What should you ask an AI visibility vendor?
Whether they can show a closed-won deal and the AI answer that started it, how many prompts, runs and engines your tier covers, and whether they report ranking position. The last one is a research-literacy test.
Continue Reading
How Does Revenue Attribution Connect Marketing to Revenue?
Run first touch and last touch side-by-side, persist origin in a read-only CRM field, and join to the billing record.
8 Best AI Visibility Tools, Ranked (2026)
Eight picks from a 50-vendor analyst landscape, each for a specific, named use case with methodology shown.
AI Visibility Tools: Market Landscape (50 Vendors)
Capability matrix across 50 vendors evaluating engine coverage, citation tracking, crawler logs, and closed-won attribution.
Schema Markup for AI: Does It Get Your Brand Cited?
Three controlled studies tested whether schema markup increases AI citations and found no direct uplift.
Sources
- SparkToro and Gumshoe.ai, “AIs are highly inconsistent when recommending brands or products”, January 2026. 600 volunteers, 2,961 runs.
- Similarweb, “The Downstream Impact of AI Visibility”, 21 June 2026. US desktop clickstream panel.
- Amsive, “Does LLM Traffic Convert Better Than Organic?”, September 2025. 54 sites, six months of GA4 data, paired t-test p = 0.794.
- Ahrefs, “Does AI Search Traffic Convert Better Than Traditional Search?”, 16 June 2025. First-party data.
- Conductor, 2026 AEO / GEO Benchmarks Report, published 13 November 2025. 13,770 domains.
- Google Search Central, “Introducing Search Generative AI performance reports in Search Console”, 3 June 2026.
- Cloudflare, “The crawl before the fall of referrals”, 1 July 2025.
- Kevin Indig, “How to measure the impact of AI search the right way”, Growth Unhinged (Kyle Poyar), 15 July 2026.