This AI visibility case study is preliminary. It covers one 90-day pilot across 20 citedby customers, over a single quarter. Read it as a first signal that still needs more time to confirm.
What happened to AI visibility score, traffic and revenue across 20 brands?
Twenty citedby customers ran a 90-day pilot. Average AI visibility score rose 26.0%, AI crawl bot traffic rose 55%, human traffic rose 6.0%, and AI-referred revenue rose 30%.
Two things skew those averages, and we would rather show them than hide them. ThriveStack's own site ran every fix the day it shipped. Its score moved from 6 to 14, a 133% jump. Take that one account out, and the typical score gain across the rest of the cohort was 20.4%.
Four smaller accounts saw revenue barely move: data-mania.com, oyubotanica.com, ordercup.com and tractionresearchlabs.com. Each was up just 1% in 90 days, even though their visibility and crawl traffic moved like everyone else's. The other sixteen averaged a 38% revenue lift. Visibility can move before revenue catches up, especially for a smaller site in a short window.
The table lists all 20 customers, their industry, and the before and after numbers we tracked. Traffic and revenue are shown as an index, with the pre-audit period set to 100. An index shows the size of the move without publishing each company's exact traffic or revenue figures. The color next to each percentage is a red, amber and green read of that single metric on its own.
| Domain | Company | Industry | AI Visibility Score | AI Crawl Bot Traffic Index | Human Traffic Index | AI-Referred Revenue Index |
|---|---|---|---|---|---|---|
| Agorapulse | Social media management software | 7 → 9 +28.6% | 100 → 166 +66% | 100 → 105 +5.2% | 100 → 140 +40% | |
| Hotel Investor Apps | Hospitality investment software | 12 → 14 +16.7% | 100 → 152 +52% | 100 → 106 +5.5% | 100 → 131 +31% | |
| AskPorter | AI property operations (PropTech) | 12 → 14 +16.7% | 100 → 162 +62% | 100 → 105 +5.2% | 100 → 137 +37% | |
| FiveX | B2B growth and RevOps software | 16 → 21 +31.2% | 100 → 170 +70% | 100 → 107 +6.6% | 100 → 132 +32% | |
| Ken | AI knowledge management | 8 → 10 +25.0% | 100 → 147 +47% | 100 → 105 +5.2% | 100 → 148 +48% | |
| PeopleKit | HR and people analytics | 11 → 13 +18.2% | 100 → 151 +51% | 100 → 106 +6.3% | 100 → 145 +45% | |
| ClickPost | Logistics and delivery tracking | 10 → 12 +20.0% | 100 → 157 +57% | 100 → 106 +6.5% | 100 → 141 +41% | |
| Timelines.ai | WhatsApp business communication | 11 → 13 +18.2% | 100 → 141 +41% | 100 → 106 +6.2% | 100 → 134 +34% | |
| RVshare | RV rental marketplace | 12 → 15 +25.0% | 100 → 148 +48% | 100 → 106 +6.3% | 100 → 134 +34% | |
| Data-Mania | Data science education and consulting | 11 → 13 +18.2% | 100 → 157 +57% | 100 → 106 +5.8% | 100 → 101.0 +1.0% | |
| LeadTagger | Sales intelligence software | 11 → 14 +27.3% | 100 → 166 +66% | 100 → 106 +6.1% | 100 → 144 +44% | |
| Traction Research Labs | GTM market research | 3 → 4 +33.3% | 100 → 155 +55% | 100 → 106 +6.2% | 100 → 101.0 +1.0% | |
| Awery Aviation Software | Aviation ERP software | 6 → 7 +16.7% | 100 → 169 +69% | 100 → 106 +6.3% | 100 → 133 +33% | |
| OrderCup | E-commerce shipping software | 7 → 8 +14.3% | 100 → 168 +68% | 100 → 107 +6.6% | 100 → 101.0 +1.0% | |
| Oyu Botanica | Botanical skincare (DTC beauty) | 8 → 9 +12.5% | 100 → 148 +48% | 100 → 106 +5.8% | 100 → 101.0 +1.0% | |
| BetterPic | AI headshot generation | 7 → 8 +14.3% | 100 → 166 +66% | 100 → 105 +5.4% | 100 → 136 +36% | |
| AvaSure | Virtual patient monitoring (health tech) | 10 → 12 +20.0% | 100 → 151 +51% | 100 → 106 +6.2% | 100 → 147 +47% | |
| FullStory | Digital experience analytics | 7 → 8 +14.3% | 100 → 141 +41% | 100 → 107 +6.6% | 100 → 136 +36% | |
| HR Cloud | HR information systems | 12 → 14 +16.7% | 100 → 146 +46% | 100 → 106 +6.3% | 100 → 135 +35% | |
| ThriveStack | Revenue intelligence platform | 6 → 14 +133.3% | 100 → 144 +44% | 100 → 106 +6.5% | 100 → 131 +31% |
Green: strong lift for that metricAmber: moderate liftRed: minimal or not yet moved
Domain, company, industry and the small site icon next to each domain are public information pulled from each company's own website. Score, traffic and revenue columns are preliminary pilot data collected by ThriveStack and have not been independently audited.
What AI search visibility factors move a brand in and out of an answer?
Eight variables move an AI answer at once, even when the prompt is identical. Search indexation, prompt quality, run frequency, engine choice and the API versus interface gap are yours to control. Persona, memory and model version are not.
Run the same prompt twice and you rarely get the same answer. In the largest public study of the question, fewer than 1 in 100 repeat runs returned the same set of brands, and fewer than 1 in 1,000 returned them in the same order.
This is why citedby's audit starts with the checklist below. The fixes aim at appearing in more runs, since rank position resets on its own anyway.
| Variable | What it changes | Observed effect | Control |
|---|---|---|---|
| 1. Search indexation (GSC / IndexNow) | Whether content enters the RAG index | Unindexed pages are 100% invisible to live search-augmented AI engines | Yes |
| 2. Prompt quality | Which brands surface | Real prompts on one intent averaged 0.081 semantic similarity | Yes |
| 3. Run frequency | Brand set and ordering | Under 1% of repeat runs return the same set of brands | Yes |
| 4. Region & locale | Which brands are known | Mention rates range roughly 18% to 50% depending on country | Partly |
| 5. AI engine | Which sources get cited | 11% source overlap between ChatGPT and Perplexity on identical prompts | Yes |
| 6. API vs. UI | Length, brands, citations | Around 24% brand overlap between the two surfaces; 406 vs. 743 words | Yes |
| 7. Persona / memory | Personalization of the answer | Account state, history and session memory shift what is returned | No |
| 8. Model version | The entire board | A rollout reshuffles rankings with no changelog you can read | No |
What does an AI brand visibility checklist fix on your site?
citedby runs a 73-point audit across four categories: technical infrastructure, content structure, entity signals and earned media. Every fix carries a before and after check, run automatically against the homepage and every key page.
- Technical infrastructure (16 fixes). Robots.txt rules that allow AI crawlers, an llms.txt file, sitemap coverage and page speed.
- Content structure (18 fixes). A direct answer in the first 40 to 60 words, question-style headings, FAQ blocks and comparison tables.
- Entity signals (19 fixes). Organization schema, FAQPage schema, a Wikidata entry and a verified Knowledge Panel.
- Earned media (20 fixes). Press coverage, guest posts, podcast mentions and third-party reviews.
Most customers in this cohort started with fewer than half the technical infrastructure items in place. An unindexed page cannot be cited, no matter how well it reads, so that category comes first.
| Category | Fixes | Examples |
|---|---|---|
| Technical infrastructure | 16 fixes | Robots.txt rules, an llms.txt file, sitemaps, page speed |
| Content structure | 18 fixes | A direct answer up top, question headings, FAQ blocks |
| Entity signals | 19 fixes | Organization schema, FAQPage schema, a Wikidata entry |
| Earned media | 20 fixes | Press coverage, guest posts, podcast mentions, reviews |
What is the F.A.C.T. diagnosis framework for AI visibility?
F.A.C.T. stands for Findable, Agent Accessible, Citable and Trustable. Every customer gets scored on all four before any fix work starts, so the team knows which category to fix first.
Findable maps to indexation: is the page in the sitemap and the llms.txt file, and can a crawler discover it at all. Agent Accessible maps to the technical infrastructure fixes: server-side rendering, clean metadata and SEO-ready markup that an AI agent can parse without running JavaScript. Citable maps to content structure: does the page name the right entities and answer the question in the first few sentences. Trustable maps to earned media: do other sites, reviews and publications independently mention the brand.
| F.A.C.T. category | What it checks | Fix category |
|---|---|---|
| Findable | Can AI crawlers and search indexes discover the page at all | Indexation, sitemap and llms.txt fixes |
| Agent Accessible | Server rendered pages with clean metadata, SEO ready | SSR, metadata and SEO-ready fixes |
| Citable | Entities named and an answer-first structure up top | Entity markup and answer-first content fixes |
| Trustable | Do third parties independently mention the brand | Earned media and citation fixes |
How do you create AI visibility prompts that target your market?
Start with buyer-intent prompts, the kind a real prospect would actually type into ChatGPT. citedby builds a starter set of 10 or more prompts across the funnel, then ranks each one by revenue potential.
Comparative prompts such as best [category] for [ICP] and competitive prompts such as [brand] vs [competitor] tend to carry the most buying intent. The panel is not fixed. It is reviewed weekly, and new prompts get added as competitors ship features or a category shifts.
Setup takes 5 to 10 minutes. After that, the weekly loop, adding or editing prompts, scanning the report, and routing gaps to the right team, takes about 30 minutes per brand. For a tactical step-by-step breakdown, explore our guide on How To Increase AI Visibility within Days.
How often should you run AI visibility tracking prompts?
Run every prompt at least seven times a day, per engine. One run carries a standard error of 0.370 on the detection rate, which a University of St. Gallen study called essentially uninformative. The error drops below 0.10 at seven runs and below 0.08 at eight.
Fewer than 1 in 100 repeat runs return the same set of brands. Answers shift with server load, with what the live index returns, and with the time of day. A single reading tells a brand almost nothing about where it actually stands.
How does an AI referral revenue platform close the loop on AI visibility?
Visibility only matters once it turns into pipeline. citedby ties every AI-referred session to first-touch revenue in the CRM, so a brand can see which engine, which prompt and which page led to a signup. Learn more in our study on Revenue Attribution for AI Search and how teams map multi-touch pipeline.
No signup required for the sample dashboard.
What AI visibility benchmark should B2B brands expect in 90 days?
Across this cohort, AI visibility score rose 26.0% on average, crawl bot traffic rose 55%, and AI-referred revenue rose 30% in 90 days. Individual results ranged from 1.0% to 48% revenue lift. Compare these metrics with our comprehensive AI Visibility Gap Report 2026 benchmarking 500+ SaaS vendors.
The next update to this report will show whether the gains held, grew or faded once the prompt panel runs for a full year across a larger cohort.
See your own AI visibility score move
citedby audits your site, tracks your prompts, and ties AI referrals to revenue.
Start free →AI visibility case study: FAQ
What is an AI visibility case study?
It is a before and after look at what happens when a brand fixes technical SEO, entity signals, content structure and prompt coverage, then re-measures how often AI engines mention and cite it. This report covers 20 citedby customers over a 90-day pilot and is preliminary.
How is AI visibility score measured?
citedby measures how often a brand appears across many repeated prompt runs on ChatGPT, Perplexity, Gemini and Google AI Mode, rather than where it ranks in any single answer. Research on non-determinism shows rank changes almost every run, while appearance frequency holds steadier and is the more honest signal.
How many prompt runs do you need before trusting an AI visibility number?
At least seven runs per prompt per engine per day, aggregated over a two to four week rolling window. A University of St. Gallen study found a single run carries a standard error of 0.370, which drops below 0.10 at seven runs.
What is the F.A.C.T. diagnosis?
F.A.C.T. stands for Findable, Agent Accessible, Citable and Trustable. citedby scores every customer on all four before recommending fixes, since a page with strong content that a crawler cannot reach still will not appear in an AI answer.
Why does citedby price by usage instead of by seat?
AI answers are non-deterministic, so tracking needs to flex with events. Usage-based pricing lets a brand run its prompt panel daily right after a fix ships, then scale back to weekly once results settle, instead of paying for a fixed cadence year round.
Is 90 days enough to call this an AI visibility benchmark?
No. This is a preliminary case study across one 90-day window and 20 customers. A secular, durable increase needs 6 to 12 more months of tracking across a larger cohort before it should be treated as a benchmark.
Sources
- ThriveStack, "AI Visibility Tracking: Why Prompts Return Different Answers": source of the non-determinism, run-frequency and "what moves a brand" data used in this report.
- ThriveStack, "AI Brand Visibility Checklist": source of the 73-point, four-category audit referenced in Tactic 1.
- ThriveStack, "How To Increase AI Visibility within Days": source of the prompt-creation guidance referenced in Tactic 3.
- citedby customer pilot cohort, n=20, preliminary internal data collected by ThriveStack, September 2026. Not independently audited.
Related Empirical Studies
Server Logs & AI Search Citation Gaps
How AI bots crawl, cache, and drop enterprise pages before citing them in answer engines.
Case StudyGoogle AI Overviews: From 2 to 100+ Citations
Eight-week empirical study on indexing funnels, Bing API integration, and GSC automation.
LLM AnalyticsAI Visibility Tracking: Non-Determinism
Why repeat prompts return different brands and why frequency is the true signal over rank.