AI Market Discovery Methodology

Overview

The LLM Authority Index AI Market Discovery Index measures how tracked brands appear in commercially relevant AI-assisted discovery, recommendation, pricing, and comparison contexts.

LLM Authority Index is the primary benchmark and source-of-record layer for the program. Industry pages publish the category-level measurements. This methodology page defines the shared research process so the same technical explanation does not need to be duplicated across every industry report.

The methodology separates four different concepts that should not be treated as interchangeable:

  1. the source query universe;
  2. the prompt-surface observation universe;
  3. brand-relevant observations; and
  4. the final qualified benchmark observations used for public metrics.

In short: collect broadly, qualify consistently, classify commercial intent, code recommendation outcomes, and publish metrics with explicit denominators.


1. Research Objective

The benchmark is designed to measure AI-mediated brand discovery and recommendation behavior within a defined category and competitive set.

It is intended to answer questions such as:

  • Which tracked brands are present in commercially relevant AI answers?
  • Which brands are valid recommendations rather than incidental mentions?
  • Which brands appear in the top three?
  • Which brands are recommended first?
  • How do these outcomes change over time?
  • Which buyer-intent contexts produce different competitive outcomes?
  • Which AI/search surfaces and external evidence sources appear around those outcomes?

The benchmark is not designed to estimate total market share or prove commercial causality.


2. Build the Category Query Universe

Each category begins with a set of high-demand queries selected using Ahrefs search-volume data as a demand proxy.

The search-demand signal is used to prioritize questions representing meaningful consumer interest. It should not be interpreted as measured prompt volume inside ChatGPT, Gemini, Copilot, Google AI Mode, Google AI Overviews, or another AI system.

The underlying research dataset should retain, where available:

  • exact query text;
  • demand signal used for selection;
  • category assignment;
  • collection period;
  • market or geography;
  • language;
  • query-cluster or normalization metadata; and
  • any distinction between a fixed recurring panel and a refreshed discovery panel.

Why use search-demand data?

AI platforms do not generally expose a complete, reliable public prompt-volume dataset comparable to traditional keyword tools. Search-demand data therefore functions as an external prioritization proxy rather than a claim about AI-platform prompt frequency.


3. Run Queries Across the AI/Search Surface Universe

The category query set is evaluated across the AI/search environments included in the active benchmark configuration.

A prompt-surface observation is:

one source query tested on one AI/search surface together with the resulting answer and associated collection metadata.

For the current industry-report implementation, a monthly research run begins with approximately 800 prompt-surface observations before qualification.

The benchmark may include environments such as:

  • ChatGPT;
  • Gemini;
  • Microsoft Copilot;
  • Google AI Mode;
  • Google AI Overviews; and
  • other AI/search surfaces defined in the active benchmark configuration.

The authoritative surface universe should be maintained centrally and versioned.

Collection Metadata

Where available, the dataset should preserve:

  • surface name;
  • model or version when exposed;
  • exact collection timestamp;
  • geography or market;
  • language;
  • signed-in or logged-out state;
  • personalization controls;
  • source query;
  • raw answer text;
  • linked or cited sources; and
  • other environment settings that could materially affect reproducibility.

4. Apply Tracked-Brand Relevance Qualification

The raw collection universe intentionally contains more observations than the public competitive benchmark requires.

The first qualification stage determines whether an observation meaningfully surfaces at least one tracked company under the benchmark's inclusion rules.

An observation may qualify when a tracked brand is:

  • recommended;
  • compared with another option;
  • evaluated in a pricing or value context;
  • explicitly discussed in another qualifying commercial context; or
  • otherwise surfaced in a manner that contributes to the consumer decision being measured.

An observation should not automatically qualify because:

  • a company appears only in an unrelated citation;
  • a domain is cited without the brand being surfaced in the relevant answer;
  • the reference is incidental to the consumer decision; or
  • the answer does not meaningfully engage with the tracked competitive set.

The purpose of this stage is research relevance, not maximizing sample size.


5. Classify Commercial Buyer Intent

Brand-relevant observations must also fit one of the benchmark's defined commercial buyer-intent classes to enter the public analysis set.

Brand Recommendation

Discovery questions in which the user asks the AI/search system to suggest brands, products, stores, or options that fit a need.

Pricing & Value

Questions in which price, affordability, discounts, value, or budget suitability materially affects the evaluation.

Multi-Brand Comparison

Questions in which two or more brands, products, or alternatives are compared directly.

Observations that pass tracked-brand relevance but do not fit the active commercial-intent framework are excluded from the final public benchmark denominator.


6. Build the Qualified Benchmark Set

Observations that survive both qualification stages form the qualified benchmark set.

This is the denominator used for public recommendation metrics unless a metric explicitly defines a different applicable observation set.

A report containing 87 qualified observations therefore does not mean only 87 AI interactions were tested. It means 87 observations from the much larger source collection satisfied the benchmark's tracked-brand relevance and commercial buyer-intent rules.

Industry reports should show the upstream collection context and the final qualified denominator near the top of the page.


7. Code Brand and Recommendation Outcomes

For each applicable observation, the benchmark records the variables needed to distinguish simple brand presence from stronger recommendation outcomes.

Coded fields can include:

  • brand presence;
  • valid recommendation status;
  • recommendation rank;
  • top-three inclusion;
  • rank-one inclusion;
  • sentiment;
  • buyer-intent class;
  • AI/search surface;
  • citation or attributable evidence source; and
  • other structured fields required by the benchmark.

The same coding logic should be applied consistently across categories and measurement periods.

For public definitions, formulas, and denominator rules, see AI Market Discovery Metric Definitions.


8. Qualified Surface Breadth

The benchmark can report qualified surface breadth: the number of tested AI/search surfaces that contributed at least one qualified benchmark observation during a measurement period.

Qualified surface breadth is a post-qualification result. It is not the number of surfaces tested.

For example, if the benchmark tests the same surface universe in two months but the final qualified observations appear across three surfaces in one month and seven in another, that does not by itself mean the collection methodology changed. It means the qualified observations were distributed across more tested environments in the second month.

This distinction should be preserved in both public reports and downstream strategic analysis.


9. Recommendation, Placement, and Presence Metrics

The benchmark distinguishes several related but non-equivalent outcomes.

Presence

A tracked brand is explicitly present in the applicable answer set.

Valid Recommendation

A tracked brand appears in a recommendation context that satisfies the benchmark's recommendation-validity rules.

Top-Three Placement

The brand appears among the first three valid recommendations.

Rank-One Placement

The brand is the first valid recommendation.

These measures should not be collapsed into a single generic "visibility" claim. A brand can be present without being recommended, or recommended without being the preferred first choice.


10. Sentiment Coding

Net sentiment summarizes the coded direction of brand sentiment in applicable observations.

The public benchmark currently interprets:

  • 1.0 as entirely positive coded sentiment;
  • 0.0 as neutral coded sentiment; and
  • negative coded mentions as reducing the score.

Sentiment should be interpreted alongside sample size, presence, and recommendation placement. A high sentiment score based on very few observations should not be treated as equivalent to a similar score supported by broad category visibility.

If the production scoring formula or numeric mapping changes, the change should be versioned and disclosed under Research Standards.


11. Source and Citation Analysis

Where an AI/search surface exposes citations, links, or attributable evidence sources, those sources can be analyzed as an additional evidence layer.

Source analysis may retain:

  • source domain or URL;
  • associated tracked brands;
  • buyer-intent class;
  • AI/search surface;
  • frequency of source appearance;
  • recommendation outcome associated with the observation; and
  • source type, such as first-party, retailer, editorial, review, comparison, or marketplace.

Source appearance is evidence about the information environment surrounding an answer. It should not automatically be treated as proof that a source caused a recommendation.


12. Modeled AI Authority Value

Modeled AI Authority Value is a separate comparative opportunity measure produced by the benchmark.

It is not measured revenue or attributable sales.

Because modeled values use demand inputs and benchmark weighting, their interpretation is governed separately under Modeled AI Authority Value.

Industry pages should label the metric as modeled wherever it appears.


13. Percentage Movement

Changes between two percentage rates are expressed in percentage points.

Example:

  • July valid recommendation coverage: 27.3%
  • August valid recommendation coverage: 44.8%
  • Movement: 17.5 percentage points

For compact tables and charts, the Index uses:

  • Up 17.5 points
  • Down 8.0 points

A percentage-point movement should not be labeled as a simple percent change because that can be mistaken for relative growth.


14. Quality Assurance

Before publication or refresh, benchmark data should pass a defined QA process.

At minimum, verify that:

  • source query counts reconcile with the configured run;
  • surface counts reconcile with the active benchmark configuration;
  • brand/entity names are normalized;
  • duplicate or malformed observations are resolved;
  • qualification rules were applied consistently;
  • public rates resolve to stored numerators and denominators;
  • recommendation ranks are internally valid;
  • sentiment values fall inside the approved coding range;
  • charts reconcile with underlying tables;
  • page dates reflect the current measurement;
  • test or illustrative data has been removed; and
  • public copy matches the current methodology version.

Where automated classification is used, ambiguous or high-impact cases should be reviewable against the underlying answer.


15. Longitudinal Benchmarking

Industry benchmarks use one evergreen URL per category.

The same page is refreshed as new measurements are available, preserving:

  • the original baseline;
  • current measurements;
  • leadership changes;
  • historical highs and lows;
  • persistent gains or declines;
  • reversals;
  • material placement changes; and
  • methodology annotations when required.

This creates a cumulative research asset rather than a collection of outdated monthly pages.


16. Methodology Changes and Versioning

Changes that could affect comparability should be documented centrally.

Examples include:

  • adding or removing an AI/search surface;
  • changing the query universe materially;
  • changing geography or collection settings;
  • changing the tracked competitive set;
  • changing buyer-intent definitions;
  • changing recommendation-validity rules;
  • changing sentiment coding;
  • changing modeled-value inputs or weights; or
  • changing the pipeline in a way that affects published values.

Where feasible, historical data should be recalculated under the updated rule. Where that is not possible, affected reports should carry a concise comparability note.

See AI Market Discovery Research Standards.


17. Relationship to CiteWorks Studio

LLM Authority Index is the primary benchmark publisher.

CiteWorks Studio may use the public benchmark and underlying research outputs as the basis for strategic interpretation, competitive analysis, and company-specific recommendations.

The editorial roles should remain distinct:

  • LLM Authority Index: what was measured, how it was measured, and what the benchmark recorded.
  • CiteWorks Studio: what the benchmark may mean strategically and what a company should investigate or do next.

This separation helps preserve provenance and reduces unnecessary duplication between the two properties.