AI Search Case Study: Why Broad AI Visibility Can Misrepresent the Sources That Matter at the Buying Stage

Case study data from car insurance and medical alerts shows why broad AI visibility can differ sharply from sources cited in buying-stage recommendations.

CASE STUDY12 minutesLast updated Sep 1, 2026By Mark Huntley, J.D.

Answer Capsule

AI systems appear to use materially different citation environments when answering broad industry questions versus commercially focused recommendation questions.

LLM Authority Index compared broad AI citation datasets with commercially focused AI recommendation studies in two unrelated categories: car insurance and medical alert systems.

The result was consistent across both industries.

In car insurance, only 3 of the top 10 citation domains in the broad dataset remained in the commercial top 10.

In medical alerts, only 2 of the top 10 citation domains overlapped.

Sources that dominated broad AI discussion—including large publishers, marketplaces, YouTube, Reddit, Wikipedia, and general informational sources—often declined sharply when prompts shifted toward questions such as which company, product, or provider a buyer should choose.

At the same time, specialist publishers, first-party providers, regulators, review authorities, and highly specific buyer-fit pages often rose dramatically.

The implication is significant:

A company can measure AI visibility across thousands of prompts and still misunderstand which sources matter when AI is helping a customer make a purchase decision.

These studies measure citation frequency, not causal influence. But they show that the citation environment changes materially with prompt intent.

The AI Search Measurement Problem

Most AI visibility platforms begin with a broad question:

Which websites and brands appear most often across AI answers about an industry?

That is useful.

But it may not answer the question companies ultimately care about:

Which sources appear when AI systems are actually recommending what a buyer should choose?

Those are not necessarily the same source universes.

A source can be highly visible when an AI system explains a market, defines a product category, discusses pricing, summarizes general information, or answers educational questions.

That does not automatically mean the same source will be prominent when the user asks:

●     What is the best provider?

●     Which company should I choose?

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

●     What is best for my particular situation?

●     Which product is safest?

●     Which company is best for seniors?

●     Which insurer is best after a DUI?

●     What is the best option for this particular buyer?

LLM Authority Index tested this distinction in two categories.

The results suggest that broad AI visibility and commercial AI recommendation visibility should be measured separately.

Broad AI visibility can look strong even when a brand is weak at the recommendation stage, which is why AI visibility and AI recommendation should be measured separately.

Study 1: Car Insurance

The car-insurance analysis compared:

Broad AI Citation Dataset

30,000 responses
210,494 citations
8,193 unique domains
50,372 unique URLs

Commercial Recommendation Dataset

15 commercially focused consensus studies
4,919 citations
332 unique domains
2,733 unique URLs

The two datasets were both about car insurance.

But they did not identify the same citation leaders.

Only 3 of the top 10 domains overlapped:

●     Progressive

●     GEICO

●     Allstate

The top-10 Jaccard similarity was just 0.18.

What Rose When Car-Insurance Prompts Became Commercial?

Several sources became dramatically more prominent when prompts moved closer to provider selection.

Source

Broad Rank

Commercial Rank

Change in Citation Share

State Farm

#21

#2

4.4×

USAA

#43

#4

13.3×

Nationwide

#23

#5

3.2×

MoneyGeek

#16

#6

2.5×

J.D. Power

#100

#7

53.2×

Erie Insurance

#122

#9

64.2×

Travelers

#55

#10

12.5×

NAIC

#128

#11

61.1×

State Farm illustrates the pattern clearly.

It represented only 1.21% of broad car-insurance citations, but 5.29% of commercial recommendation citations.

Its citation share increased 4.4×, while its domain rank moved from #21 to #2.

J.D. Power and the National Association of Insurance Commissioners were even more striking.

Both were relatively minor sources in broad discussion but became major sources around commercially focused answers.

That suggests that the evidence AI systems surface around a buying decision can differ substantially from the information they surface when simply discussing the category.

J.D. Power citation share

A brand appearing in an AI answer does not necessarily mean it is being endorsed, reinforcing the difference between an AI mention and an actual recommendation.

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

What Fell When Car-Insurance Prompts Became Commercial?

Some of the strongest broad-AI sources went in the opposite direction.

Source

Broad Rank

Commercial Rank

Commercial vs. Broad Share

NerdWallet

#1

#13

0.34×

The Zebra

#2

#25

0.15×

Insurify

#3

#16

0.31×

U.S. News

#4

#38

0.08×

YouTube

#9

#53

0.03×

Reddit

#12

#32

0.21×

This doesn't mean these are weak sources.

It means their prominence depended heavily on what the AI system was being asked.

A company looking only at the broad dataset might conclude that quote marketplaces and large publishers define the citation environment.

A company looking at buyer-selection prompts would see a substantially different picture.

Study 2: Medical Alert Systems

The medical-alert analysis produced an even larger difference.

LLM Authority Index compared:

Broad AI Citation Dataset

26,773 responses
151,601 citations

Commercial Recommendation Dataset

15 Medical Alert consensus studies
4,534 citations

Only 2 of the top 10 citation domains appeared in both datasets:

●     NCOA.org

●     Google.com

Top-10 Jaccard similarity was just 0.11.

Even expanding the comparison did not eliminate the difference:

●     Top 25 overlap: 5 of 25

●     Top 50 overlap: 15 of 50

Commercial recommendation citations were also considerably more concentrated.

The ten largest domains accounted for:

23.8% of broad citations

versus

47.6% of commercial recommendation citations.

 23.8% of broad citations but 47.6% of commercial citations

Explore our citation architecture gap study to see how differences in source coverage can influence which brands AI platforms cite and recommend.

What Rose in Commercial Medical-Alert Recommendations?

The sources gaining prominence again tended to be much more closely associated with buyer evaluation and specific products.

Source

Broad Rank

Commercial Rank

Change in Citation Share

SeniorLiving.org

#12

#1

6.8×

NCOA.org

#9

#2

4.1×

SafeWise

#49

#3

18.3×

Bay Alarm Medical

#35

#4

9.9×

The Senior List

#28

#6

7.1×

Medical Guardian

#24

#7

6.1×

MobileHelp

#147

#9

58.6×

Lively

#103

#10

31.2×

Caring.com

#91

#15

16.0×

NCOA provides a particularly clean example.

It represented 1.45% of citations in broad medical-alert AI discussion, but 5.91% of commercial recommendation citations.

That is a 4.1× higher citation share, with its ranking moving from #9 to #2.

medical alerts citation share

This gap is explored further in the commercial AI citation blindspot, where broad citation visibility can obscure the sources that dominate buyer-stage recommendations.

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

What Fell in Medical Alerts?

General discussion sources collapsed even more dramatically.

Source

Broad Rank

Commercial Rank

Commercial vs. Broad Share

YouTube

#1

#49

0.01×

Reddit

#6

#48

0.03×

Wikipedia

#11

#50

0.02×

NIH

#22

#50

0.03×

Mayo Clinic

#38

Absent

CDC

#80

Absent

YouTube was the #1 citation domain in the broad dataset, representing 4.10% of citations.

In the commercially focused dataset, it fell to #49 and just 0.04%.

Reddit dropped from #6 to #48.

Wikipedia dropped from #11 to #50.

Meanwhile, specialist comparison sites, senior-focused publishers, and medical-alert providers became substantially more prominent.

Two Industries, the Same Structural Finding

Car insurance and medical alert systems have very different products, audiences, economics, regulatory environments, publishers, and buyer journeys.

Yet the underlying pattern was remarkably similar.

Finding

Car Insurance

Medical Alerts

Top-10 domain overlap

3/10

2/10

Broad citation leaders remained dominant commercially?

Mostly no

Mostly no

Specialist/buyer-fit sources rose?

Yes

Yes

Generic/discussion sources fell?

Yes

Yes

Commercial citation environment differed materially?

Yes

Yes

broad vs. commercial citation source overlap

This does not prove that every industry will behave identically.

It does, however, establish a strong reason to test the phenomenon in any market where consumers use AI systems to compare products, evaluate companies, or select providers.

The same pattern appears at the brand level in Life Alert’s AI visibility without recommendation, where visibility did not translate into recommendation strength.

Why This Could Matter Across Industries

Consider how the same distinction could affect different markets.

A broad AI dataset about mortgages might identify the websites most frequently cited when explaining interest rates, home buying, or lending.

That may not identify the sources used when someone asks:

“Which HELOC lender is best for someone with excellent credit?”

A broad dataset about software might identify technology publishers and large review sites.

That may not identify the evidence used when someone asks:

“Which project management platform is best for a 200-person engineering team?”

A broad dataset about tax relief might surface government information, news, forums, and financial education sites.

That may differ from the sources used when someone asks:

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

“Which tax relief company is most trustworthy for someone who owes the IRS $80,000?”

A broad dataset about home appliances may contain reviews, YouTube videos, Reddit threads, manufacturer information, and shopping pages.

The source mix may change again when a buyer asks:

“What is the best vacuum for a house with two shedding dogs and hardwood floors?”

The commercial question is more constrained.

So the relevant evidence set may also become more constrained.

That is the hypothesis companies should test.

The Bigger Issue: Prompt Universe Determines What You Measure

An AI visibility index is only as meaningful as its prompt universe.

If thousands of informational, navigational, loosely related, or general discussion prompts are blended together with high-intent buying prompts, the resulting citation rankings may accurately describe general AI discussion while obscuring AI-mediated purchasing decisions.

That distinction is fundamental.

Consider two metrics:

Broad Citation Share

How frequently a domain is cited across a broad universe of industry-related AI responses.

Commercial Citation Share

How frequently the domain is cited across carefully defined prompts where users are evaluating, comparing, or selecting products or providers.

Neither is inherently "correct" or "incorrect."

They answer different questions.

The problem occurs when broad visibility is presented as if it automatically represents buyer-choice visibility.

Why URL-Level Analysis Matters Too

The difference can exist even within the same domain.

The MoneyGeek analysis illustrates this.

MoneyGeek appeared in both car-insurance universes, but AI systems frequently cited completely different pages.

Broad AI cited 737 unique MoneyGeek URLs.

Commercial studies cited 76.

Only 49 URLs appeared in both datasets, giving the page sets a Jaccard similarity of only 0.06.

Commercial prompts disproportionately surfaced pages addressing specific buyer jobs such as:

●     SR-22 insurance

●     modified vehicles

●     home and auto bundles

●     senior drivers

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

●     young adults

Generic "best companies," cheap-insurance, and state pages were far more prominent in broad AI discussion.

So even saying:

“MoneyGeek is important to AI.”

is incomplete.

The more actionable question is:

Which MoneyGeek page is important for which buyer question?

That principle could apply to virtually any sufficiently large website.

This Changes How Companies Should Think About AI Search Optimization

If these patterns continue to appear in other industries, AI Search strategy should move beyond simply identifying the most-cited domains.

Companies may need to map at least four dimensions:

1. Prompt Intent

What was the user actually trying to accomplish?

Informational prompts and purchase-selection prompts should not automatically be treated as equivalent.

2. Citation Source

Which domain appeared in the answer?

This identifies the external evidence environment.

3. Cited Page

Which specific URL appeared?

Domain authority alone may conceal highly specialized page-level citation behavior.

4. Recommendation Outcome

Which company did the AI system ultimately recommend, rank, compare, or exclude?

Citation visibility is useful.

But the commercial question is whether that evidence environment corresponds with buyer choice.

What This Means for Brands

For a company, there could be a meaningful difference between:

“The websites that talk about our market.”

and:

“The websites AI systems use when deciding whether to recommend us.”

That creates several possible blind spots.

A company could have strong broad AI visibility but weak recommendation visibility.

A competitor could appear relatively minor in generic AI reporting while dominating specific high-value buyer prompts.

A publisher could appear highly authoritative at the domain level while having almost no presence around the exact commercial questions that matter.

And a small set of highly specific pages could matter disproportionately within a particular buyer-intent cluster.

This is why AI Search measurement should increasingly move from:

visibility

to:

recommendation quality and buyer-choice intelligence.

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

What This Means for Publishers

Publishers face the same issue.

The data suggests that producing more broadly popular category content does not necessarily maximize citation presence within commercial AI answers.

Content aligned to a specific buyer job may behave very differently.

Examples include:

●     best provider for a particular customer type,

●     best product for a particular use case,

●     provider comparisons,

●     safety or trust evaluations,

●     pricing decisions,

●     specialist requirements,

●     post-event situations,

●     regulatory or eligibility circumstances.

The important question may increasingly become:

What decision does this page help an AI system make?

rather than simply:

How much search volume does this keyword have?

This is also why off-intent AI visibility can create a misleading picture of performance when prompts do not reflect the buyer questions that actually matter.

What This Study Does Not Prove

These studies do not demonstrate that a citation causes an AI system to recommend a company.

They do not determine:

●     model weighting,

●     causal influence,

●     training importance,

●     retrieval importance,

●     conversion rates,

●     revenue attribution,

●     or whether changing a cited source will change an answer.

This is a citation-frequency analysis.

The finding should therefore be stated carefully:

Across the car-insurance and medical-alert datasets analyzed, the domains and URLs most frequently cited in broad industry AI discussion differed materially from those cited in commercially focused recommendation prompts.

That is the evidence.

The broader cross-industry implication is a hypothesis supported by those observations, not yet a universal law.

Implications for AI Visibility Platforms and GEO Measurement

This finding also raises an important question about AI visibility software.

A platform may process tens of thousands or millions of prompts.

But dataset size alone does not guarantee commercial relevance.

A smaller dataset of carefully selected decision-stage prompts may answer a completely different business question than a massive broad-topic dataset.

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

That means companies evaluating AI Search measurement should ask:

●     What prompts are included?

●     How are those prompts classified?

●     How much of the dataset represents buyer intent?

●     Are commercial and informational prompts separated?

●     Is performance measured at the domain level or URL level?

●     Are citations connected to the recommendation produced?

●     Are brand mentions distinguished from actual recommendations?

●     Are competitors analyzed within the same buyer-intent context?

The key issue is not simply how much data a system has.

It is what question the data is capable of answering.

Frequently Asked Questions

1. Do AI systems use different sources for informational and commercial questions?

In the two industries analyzed here, citation patterns changed substantially when prompts shifted from broad industry discussion toward commercially focused recommendation questions.

2. Does broad AI citation share measure purchasing influence?

Not by itself. Broad citation share measures frequency within the prompt universe being analyzed. It does not demonstrate that those sources influence purchasing recommendations.

3. What is commercial citation share?

Commercial citation share measures how frequently a source appears within a defined universe of high-intent comparison, recommendation, or buyer-selection prompts.

4. Why does prompt intent matter for AI Search?

Because the citation environment can change according to what the user asks. A broad informational question and a specific buyer-selection question may cause AI systems to draw on very different evidence.

5. Should informational prompts be excluded from AI visibility reporting?

Not necessarily. Informational prompts remain important for awareness and discovery. They should simply not automatically be blended with commercial prompts and interpreted as the same KPI.

6. Do domain rankings tell the whole story?

No. The car-insurance data showed substantial URL-level differences even within the same publisher. Domain visibility should therefore be supplemented with page-level and prompt-level analysis.

7. Does being cited mean a source influenced an AI recommendation?

Not necessarily. Citation frequency and causal influence are different concepts. The studies reported here measure citation frequency.

8. Could this pattern apply outside car insurance and medical alerts?

Potentially. The studies provide a reason to test the same distinction in other industries, particularly categories where users ask AI systems to compare products, evaluate companies, or make purchase decisions. Additional industry studies are required before claiming the pattern is universal.

Methodology

This analysis used two independent industry comparisons.

Car Insurance

Broad dataset:

30,000 AI responses
210,494 citations

Commercial dataset:

15 commercially focused consensus studies
4,919 citations

Medical Alert Systems

Broad dataset:

26,773 AI responses
151,601 citations

Commercial dataset:

15 commercially focused consensus studies
4,534 citations

Domains were compared using citation frequency, citation share, ranking position, top-domain overlap, and where appropriate URL-level overlap.

The analysis evaluates frequency, not influence or causation.

Final Takeaway

AI Search has at least two different visibility questions.

The first is:

Who appears when AI talks about an industry?

The second is:

Who appears when AI helps someone choose what to buy?

The car-insurance and medical-alert datasets show that those questions can produce dramatically different answers.

That distinction could have implications across any industry where AI systems increasingly participate in comparison, evaluation, shortlisting, and purchase decisions.

For brands, publishers, SEO teams, GEO teams, and AI visibility platforms, the takeaway is simple:

Do not assume the sources dominating general AI discussion are the same sources dominating AI recommendations.

Measure both.

And when commercial outcomes matter, measure the citation environment around the prompts where the buyer is actually choosing.

 

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.