What Would Prove the AI Commercial Momentum Hypothesis Wrong?
A precommitted falsification framework for AI commercial momentum, outlining the tests that could weaken or disprove it as an investor signal.
On this page
- 01Answer Capsule
- 02The Precommitted Failure Tests
- 03Questions This Section Answers
- 04What Would Actually Disprove the Hypothesis?
- 05Why Publish Failure Conditions Before the Outcome Is Known?
- 06Falsification Test 1: AI Movement Must Come Before the Outcome
- 07Falsification Test 2: The Signal Must Predict Something Outside AI Search
- 08Falsification Test 3: AI Variables Must Add Information Beyond Ordinary Baselines
- 09Falsification Test 4: The Relationship Must Survive Out-of-Sample Testing
- 10Falsification Test 5: One Platform Cannot Quietly Become the Whole Signal
- 11Falsification Test 6: One-Month Volatility Cannot Be Mistaken for Momentum
- 12Falsification Test 7: The Result Must Survive Reasonable Prompt Tests
Research status: Prospective falsification framework for exploratory longitudinal research. AI recommendation momentum has not been validated as a predictor of revenue, earnings, analyst revisions, valuation, or stock returns.
Initial observation window: July through September 2026
Initial public-company panel: 25 mapped public parents
Current methodology version: V0
Answer Capsule
The AI Commercial Momentum Hypothesis would be weakened or disproved if changes in AI recommendation visibility fail to precede economically meaningful outcomes, fail to add information beyond ordinary financial and digital indicators, disappear out of sample, reverse too quickly to be useful, or prove to be artifacts of platform volatility, prompt selection, entity mapping, or measurement choices.
LLM Authority Index is publishing these failure conditions before the downstream financial outcomes are known. That matters because a useful investment signal must survive prospective testing, not merely explain the past after the result is visible.
The core hypothesis is that persistent changes in how often companies are recommended by major AI systems for unbranded, commercially relevant questions may contain information that precedes changes in consumer consideration and possibly later commercial or financial outcomes. The hypothesis does not require AI recommendation momentum to predict every company, every sector, or every time horizon. But it does require the signal to demonstrate repeatable, incremental information value somewhere clearly defined in advance.
If future results show no stable relationship with downstream behavior after proper controls, then the correct conclusion will be that AI recommendation momentum is an interesting AI-search measurement but not a validated investor signal.
For the original hypothesis, see Can AI Search Signal Future Revenue Growth? The AI Commercial Momentum Hypothesis. For the exact V0 construction rules, see How We Measure AI Commercial Momentum. The dated company observations will be maintained in the AI Investor Signal Tracker.
The Precommitted Failure Tests
Test | What would count against the hypothesis | What would still leave the hypothesis alive |
|---|---|---|
Time ordering | AI movement occurs after, rather than before, commercial or financial changes | AI movement consistently precedes at least one meaningful downstream outcome |
Revenue relationship | No relationship with later revenue growth, revenue surprise, or revenue-estimate revisions | A stable relationship appears in defined sectors or horizons |
Incremental information | AI variables add no explanatory or predictive value beyond traditional baselines | AI variables improve out-of-sample prediction or ranking beyond baseline variables |
Out-of-sample stability | Results appear in the discovery period but vanish in later months or new companies | Direction and effect persist in walk-forward tests |
Cross-platform robustness | Apparent signal is driven mainly by one AI platform and changes sign elsewhere | Multi-platform breadth or a clearly defined platform-specific effect is repeatable |
Persistence | One-month changes reverse so quickly that they carry no durable information | Multi-month persistence improves signal quality |
Prompt robustness | Results disappear under reasonable fixed-prompt or prompt-subset tests | Results remain directionally similar across predeclared prompt subsets |
Entity mapping | Parent-company effects are actually narrow product or brand artifacts unrelated to parent economics | Segment-weighted mapping preserves the relationship |
Cleaning sensitivity | Signal direction depends materially on duplicate handling, capture selection, or extraction rules | Direction survives reasonable cleaning alternatives |
Economic significance | Statistical association is too small or unstable to matter economically | Effect size is large enough to improve real forecasting or due-diligence decisions |
Market relevance | AI movement contains no information about analyst revisions, surprises, or future returns after risk controls | Defined relationships survive risk, sector, and valuation controls |
Replication | Findings cannot be replicated in new sectors, time periods, or independent datasets | Results replicate under separately collected data |
Want the full Authority Index
The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.
Questions This Section Answers
- What exactly would disprove the AI Commercial Momentum Hypothesis?
- Does the hypothesis fail if AI visibility does not predict stock returns?
- Could the hypothesis still be useful if it works only in certain sectors?
What Would Actually Disprove the Hypothesis?
The strongest version of the hypothesis would be disproved if AI recommendation momentum repeatedly fails three tests at the same time:
- No forward relationship: recommendation gains and losses do not reliably precede any meaningful commercial or financial outcome.
- No incremental value: adding AI recommendation variables to a reasonable baseline model does not improve prediction, ranking, or explanatory power out of sample.
- No stable domain of usefulness: the relationship does not become useful even after separating sectors, platform breadth, persistence, prompt intent, or economically relevant business segments using rules defined before evaluating the result.
If all three conditions hold after enough data accumulates, then the responsible conclusion is not that the model needs another score or another label. The responsible conclusion is that the investment thesis failed.
A failure to predict stock returns alone would not automatically disprove the narrower commercial hypothesis. Stock prices incorporate many variables unrelated to AI-mediated consumer demand, including valuation, rates, capital structure, macro conditions, regulatory events, earnings quality, guidance, and expectations already embedded in price.
The research ladder therefore moves from easier and more directly connected outcomes to harder ones:
AI recommendation momentum → brand search and digital engagement → commercial demand indicators → revenue expectations → reported revenue and earnings → analyst revisions → sector-relative stock returns
If AI momentum predicts branded search but does not predict revenue, that would narrow the hypothesis. If it predicts revenue revisions but not stock returns, that would also narrow the hypothesis. If it predicts nothing beyond AI-search behavior itself, the investor thesis would fail.
Likewise, the hypothesis could remain useful if the relationship exists only in certain sectors, provided those sectors are identified through prospective testing rather than selected after the fact. Consumer-facing financial services, insurance, healthcare, lending, e-commerce, travel, software, and other recommendation-sensitive categories may behave differently from commodity producers, regulated utilities, or businesses whose revenue is only weakly connected to consumer AI discovery.
That is why later validation will distinguish a universal hypothesis from a sector-conditional hypothesis rather than treating all public companies as equally exposed to AI-mediated consideration.
Want the full Authority Index
The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.
Why Publish Failure Conditions Before the Outcome Is Known?
The central risk in any new alternative-data research program is retrospective storytelling.
If researchers wait until they know which companies grew, which stocks outperformed, and which analysts revised estimates, it becomes easy to search backward for an AI metric that happened to move in the same direction.
That is not the standard we want for this research.
The initial LLM Authority Index series therefore freezes four things publicly:
- The hypothesis. AI recommendation momentum may precede changes in economically meaningful outcomes.
- The measurement system. The V0 rules specify matched prompt-platform cells, failure handling, duplicate handling, parent-company mapping, uncertainty, platform breadth, candidate thresholds, and sensitivity analysis.
- The initial signals. The complete first 25-company panel is already published in Initial Findings From 25 Public Companies.
- The failure conditions. This article defines what future evidence would count against the thesis.
That creates a dated record.
If the research later works, readers can verify that the claimed relationship was not invented after the outcome. If it fails, the same archive will show that failure clearly.
This is also why the project will preserve older monthly publications rather than silently rewriting them when later data changes the interpretation.
Falsification Test 1: AI Movement Must Come Before the Outcome
Questions This Section Answers
- Why is time ordering essential?
- What if AI recommendations simply reflect news that investors already know?
- How will future tests avoid look-ahead bias?
A leading indicator must lead.
If AI recommendation visibility changes only after revenue acceleration, analyst upgrades, media coverage, product launches, or stock-price moves are already visible, then it may be descriptive without being predictive.
The first validation requirement is therefore strict time ordering.
The frozen AI signal at time T must be compared with outcomes measured at T+1, T+2, or another predeclared forward horizon. Financial variables that became known after the AI observation date cannot be used to construct the original signal.
This distinction is especially important because AI systems themselves may ingest public information about company momentum. A company experiencing strong growth may receive more press coverage, more reviews, more publisher attention, more user discussion, and more web content. AI recommendations could then rise because the business was already improving.
That would still be commercially interesting, but it would not establish a leading investor signal.
The planned revenue-growth validation study will therefore preserve strict temporal ordering. The backtesting framework will use walk-forward or otherwise time-respecting tests rather than random train-test splits that leak future information backward.
Want the full Authority Index
The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.
Falsification Test 2: The Signal Must Predict Something Outside AI Search
A measurement can be internally consistent without being economically useful.
Suppose a company gains recommendation coverage across ChatGPT, Gemini, Google AI Mode, Google AI Overviews, Microsoft Copilot, and Perplexity. If that movement has no relationship with anything outside those AI systems, then the investor interpretation is unsupported.
The first downstream outcomes should therefore be relatively close to the hypothesized mechanism:
- branded search demand;
- direct website traffic;
- app traffic or engagement where observable;
- retailer or marketplace navigation;
- conversion or acquisition metrics where obtainable;
- AI-referred traffic and revenue where measurable.
Only after establishing a relationship with nearer commercial outcomes should the research move confidently toward revenue, earnings, or stock-price interpretation.
A 2026 observational study, From Prompt to Purchase: How AI Brand Recommendations Move Consumers on the Open Web, found that users with no recent observed brand engagement showed higher same-name Google search and brand-site visitation after AI brand recommendations. The authors explicitly describe the design as observational and do not observe transactions. That makes the paper relevant to the mechanism, not proof of the LLMAI investor hypothesis. See Iannelli and Ai, 2026.
Likewise, a 2026 Marketing Science study analyzed 12 months of first-party e-commerce data from 973 websites, including more than 50,000 ChatGPT-referred transactions. The study demonstrates that organic LLM traffic can be measured against conversion and revenue outcomes, but it does not test whether recommendation momentum predicts future public-company financial results. See Kaiser and Schulze, 2026.
Those studies make the downstream pathway testable. They do not make it proven.
Falsification Test 3: AI Variables Must Add Information Beyond Ordinary Baselines
This may be the most important test in the entire project.
If investors can obtain the same information more accurately from traditional variables, AI recommendation data may be redundant.
A proper baseline should eventually include variables such as:
- prior revenue growth;
- analyst revenue and EPS estimates;
- recent estimate revisions;
- earnings surprise history;
- valuation multiples;
- sector and industry controls;
- market capitalization;
- stock-price momentum;
- branded search trends;
- website or app traffic where available;
- other relevant demand indicators.
The research question is not merely:
Does AI recommendation momentum correlate with future growth?
It is:
Does AI recommendation momentum contribute incremental information after accounting for the information investors already have?
If a baseline model predicts future revenue revisions with a certain level of error and the addition of AI variables does not improve that error out of sample, then the AI variables have not demonstrated incremental forecasting value.
Similarly, if the cross-sectional rank correlation between AI momentum and later financial outcomes is indistinguishable from noise after ordinary controls, the investment thesis weakens substantially.
The signal should therefore be evaluated with measures such as:
- change in out-of-sample R-squared;
- change in mean absolute error;
- directional accuracy;
- area under the curve for directional classification where appropriate;
- information coefficient or rank correlation;
- sector-neutral top-minus-bottom portfolio spreads;
- turnover and implementation costs for market-return tests.
No single statistic will decide the thesis. The important question is whether the improvement is stable, economically meaningful, and repeatable.
Want the full Authority Index
The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.
Falsification Test 4: The Relationship Must Survive Out-of-Sample Testing
A signal discovered and evaluated on the same observations is vulnerable to overfitting.
The current V0 panel should therefore be treated as a starting cohort, not a final proof sample.
The research program should progressively test:
- later months for the same companies;
- newly added companies not used to tune the first rules;
- additional sectors;
- different economic environments;
- different AI platform versions;
- fixed prompt universes that were defined before outcome observation.
A result that looks strong from July through September 2026 but disappears from October onward would be weak evidence.
A more convincing pattern would survive repeated walk-forward testing in which each period's AI signal is frozen before the next period's commercial or financial outcome is observed.
This is why monthly publication history is part of the methodology rather than merely a content strategy.
The AI Investor Signal Tracker is intended to preserve those dated observations.
Falsification Test 5: One Platform Cannot Quietly Become the Whole Signal
AI platforms do not produce identical recommendations.
The initial V0 method separately tracks six platform families:
- ChatGPT;
- Gemini;
- Google AI Mode;
- Google AI Overviews;
- Microsoft Copilot;
- Perplexity.
A company can gain sharply on one platform while declining on another.
If the apparent investor signal is repeatedly driven by one platform while the others move randomly or in the opposite direction, then a broad "AI Commercial Momentum" interpretation would be too strong.
There are two possible outcomes:
- Broad portability exists. Cross-platform movement contains the stronger signal.
- Platform-specific effects exist. One or two systems prove economically relevant while others do not.
Either result can be useful, but they imply different models.
The error would be to discover later that only one platform works and then retroactively describe the original six-platform signal as if that had always been the hypothesis.
The dedicated cross-platform AI visibility study will test this directly. It also connects with the broader LLM Authority Index Persistence-Portability Gap, which distinguishes persistence over time from portability across AI systems.
Want the full Authority Index
The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.
Falsification Test 6: One-Month Volatility Cannot Be Mistaken for Momentum
AI responses are stochastic and model behavior changes.
A one-month spike may reflect:
- model updates;
- temporary source retrieval differences;
- changing web freshness;
- prompt interpretation variance;
- transient news coverage;
- random response variation;
- changing platform search behavior.
If positive and negative signals reverse frequently from one month to the next, then the economic usefulness of raw monthly movement may be low.
That would not necessarily mean AI visibility is useless. It could mean the useful variable is persistence, not first difference.
Possible future refinements that should be tested prospectively include:
- two-month persistence;
- three-month persistence;
- rolling recommendation coverage;
- exponential weighting of recent observations;
- percentage of platform families maintaining the same direction;
- duration above or below a company-specific baseline.
If only persistent multi-month movement predicts downstream outcomes, the original one-period signal would need to be narrowed accordingly.
If even persistent movement fails, the hypothesis weakens further.
Falsification Test 7: The Result Must Survive Reasonable Prompt Tests
Prompt design is a major source of potential bias in AI visibility measurement.
An investor signal cannot depend on choosing prompts that happen to favor the desired company or sector.
The current methodology reduces this problem by comparing the same normalized prompt on the same platform family across periods. That protects the time-series comparison from simple prompt-mix changes.
But future validation should go further.
Reasonable robustness tests include:
- fixed high-intent prompt sets;
- prompt subsets by buyer stage;
- prompt subsets by commercial intent;
- prompt subsets by product or service line;
- exclusion of prompts with unusually narrow brand implications;
- balanced prompt weights across sectors;
- sensitivity to prompt-cluster definitions.
If the signal changes direction under small, reasonable prompt-set changes, then the result may reflect sampling rather than company momentum.
If the direction survives predeclared prompt subsets, confidence increases.
This is also why the methodology should avoid optimizing prompt sets against known future financial performance. Prompt selection must remain an AI-measurement decision, not a hidden outcome-fitting step.
Want the full Authority Index
The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.
Falsification Test 8: Parent-Company Mapping Must Match Economic Exposure
The initial public-company panel maps brands, products, and entity variants to listed parents.
That is necessary for investor analysis, but it creates a clear risk.
Examples in the current V0 panel include:
- Marcus by Goldman Sachs mapped to Goldman Sachs Group;
- Coinbase Wallet mapped to Coinbase Global;
- CVS Pharmacy mapped to CVS Health;
- Labcorp OnDemand mapped to Labcorp;
- UFB Direct mapped to Axos Financial;
- UnitedHealthcare consumer entities mapped to UnitedHealth Group.
If the tracked product represents only a small part of the parent's economics, a large AI movement may have little effect on consolidated revenue.
Future validation therefore needs segment-aware weighting wherever possible.
A company-level signal should be weakened or rejected when:
- the tracked entity has low revenue relevance;
- the mapped product is not strategically important;
- the parent earns most revenue through unrelated channels;
- recommendation changes occur in a legacy brand during rebranding;
- the public parent has substantial non-consumer businesses unrelated to the prompt universe.
The Happen/LendingClub row is a clear example of why this matters. A decline in recommendations for the legacy LendingClub entity during a rebrand may measure brand migration or AI staleness rather than deteriorating demand.
A stronger future system should use a governed entity master and, where data permits, business-segment revenue weights.
Falsification Test 9: The Signal Must Survive Cleaning and Capture Sensitivity
A valid research signal should not depend on one fragile cleaning decision.
The V0 methodology already includes two sensitivity checks:
- No-dedupe sensitivity: recalculate without collapsing cross-vertical response-identical duplicates.
- Capture-average sensitivity: average multiple distinct retained response captures within the same prompt-platform-month cell before comparing months.
In the initial 25-public-parent panel, the median difference between the clean signal and the no-dedupe sensitivity result was 0.0 percentage points. The capture-average sensitivity difference also had a median effectively equal to zero.
Those results are encouraging for measurement stability, but they do not validate financial usefulness.
Future versions should continue to test whether:
- company direction changes under cleaning alternatives;
- candidate classification changes under cleaning alternatives;
- effect size materially depends on response consolidation;
- results depend on excluding a small number of high-failure datasets;
- results depend on one normalization rule.
If the strongest apparent financial relationships disappear under reasonable data-cleaning alternatives, confidence should fall.
Want the full Authority Index
The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.
Falsification Test 10: Statistical Significance Is Not Enough
Even if an AI signal is statistically associated with future outcomes, it may be economically trivial.
For example, a model might detect a small average relationship between recommendation momentum and revenue revisions but produce too little improvement to matter for investors, corporate strategists, or forecasters.
The research should therefore evaluate both statistical and economic significance.
Questions include:
- Does the signal materially reduce forecast error?
- Does it improve the ranking of companies by future commercial outcome?
- Are top and bottom signal groups meaningfully different after sector adjustment?
- Does any stock-return spread survive transaction costs and turnover?
- Is the effect large enough to change a reasonable diligence conclusion?
- Does it remain useful after traditional data are included?
A tiny but statistically significant result would not justify strong investor language.
Falsification Test 11: Market Predictions Need a Higher Standard Than Revenue Predictions
Predicting company operations and predicting stock returns are different problems.
A company can experience improving commercial momentum while its stock falls because the improvement was already priced in, valuation was excessive, rates changed, margins deteriorated, guidance disappointed, or another business segment weakened.
Similarly, a company can experience weak AI recommendation momentum while its stock rises for reasons unrelated to consumer AI discovery.
This is why the research sequence separates:
- commercial performance;
- analyst expectations;
- market expectations;
- valuation;
- stock returns.
The later concept of AI/Market Divergence should only be tested after the commercial relationship is established. See AI Visibility Market Divergence: Can AI Recommendation Momentum Reveal Information Not Yet Reflected in Investor Expectations?.
Likewise, a future comparison of AI Recommendation Share vs. Market Share should not assume that recommendation share mechanically converts into purchase share or enterprise value.
That distinction is reinforced by recent prior art. On September 23, 2026, AIVO withdrew previously published Grüns LLM Equity Valuation figures after concluding that its original formula overstated AI-reachable revenue and compared quantities measured in incompatible units. Its revised framework explicitly calls for longitudinal testing of whether recommendation-share changes precede AI-referred revenue before relying on the method for pricing. See AIVO's correction.
That is exactly the type of methodological failure this LLMAI series is designed to avoid.
Want the full Authority Index
The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.
Falsification Test 12: Replication Must Eventually Extend Beyond LLMAI's Original Corpus
A signal is more credible when independent data collection produces similar results.
The current investor series begins with LLM Authority Index's preserved July through September 2026 corpus. That is appropriate for developing and prospectively freezing the method, but it is not sufficient for final validation.
Longer-term replication should include:
- independently collected prompt panels;
- new sectors;
- companies not used in V0 development;
- future model versions;
- different geographic markets where appropriate;
- different prompt authors or prompt-construction procedures;
- third-party commercial outcome datasets.
If the relationship exists only inside one proprietary archive and cannot be reproduced elsewhere, the evidence would remain limited.
Three Possible Outcomes of the Research Program
The project does not have only two possible endings.
Outcome 1: Strong validation
AI recommendation momentum consistently precedes one or more downstream commercial outcomes, adds incremental information to traditional baselines, survives out-of-sample testing, and remains robust across reasonable measurement choices.
Under that outcome, the research could progress from AI Commercial Momentum toward a validated AI Commercial Leading Indicator.
Only after additional market testing would concepts such as AI/market divergence or valuation classification become appropriate to evaluate.
Outcome 2: Conditional usefulness
The signal works only under specific conditions, such as:
- consumer-facing sectors;
- high-intent prompt clusters;
- persistent multi-month movement;
- broad cross-platform movement;
- certain AI platforms;
- companies with high digital acquisition exposure;
- products where AI materially participates in consideration.
That would not be failure. It would mean the original broad hypothesis should be narrowed to the conditions supported by evidence.
Outcome 3: Disconfirmation
Recommendation momentum does not consistently precede meaningful business outcomes, adds no incremental information beyond traditional variables, fails out of sample, or proves too unstable to use.
Under that outcome, LLM Authority Index should continue to describe recommendation momentum as an AI-search measurement but should stop presenting it as a candidate financial leading indicator.
That is a valid research result.
Want the full Authority Index
The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.
Questions This Section Answers
- Can the methodology change over time without invalidating the study?
- How should failed hypotheses or revised rules be published?
- What prevents the research from moving the goalposts?
How Methodology Changes Will Be Handled
The methodology can improve, but historical versions must remain visible.
If future research shows that V0 should be changed, the update should be versioned rather than silently replacing the old rules.
For example:
- V0: initial matched recommendation-coverage methodology;
- V1: hypothetical future version adding validated persistence weighting;
- V2: hypothetical future version adding segment exposure weighting.
A new version should state:
- what changed;
- why it changed;
- which historical results were produced under the old method;
- whether prior classifications would have changed;
- whether the new rule was defined before evaluating the next outcome period.
Historical monthly articles should remain available.
This approach allows methodology improvement without erasing the prospective record.
The canonical technical reference for the current version remains How We Measure AI Commercial Momentum.
What We Will Test First
The next validation layer should begin with outcomes closest to the hypothesized commercial mechanism rather than jumping directly to stock performance.
Near-term validation
Potential 30-day or similar forward outcomes:
- branded search direction;
- web or app traffic direction;
- observable AI-referred traffic;
- analyst estimate revisions where timing allows.
Intermediate validation
Potential 90-day or quarterly outcomes:
- revenue growth;
- revenue surprise;
- EPS surprise;
- company guidance changes;
- analyst revenue and EPS revision direction.
Longer-horizon validation
Potential 180-day and longer outcomes:
- sector-relative stock returns;
- risk-adjusted returns;
- valuation changes;
- persistence of analyst revisions;
- multi-quarter revenue trajectories.
The exact research design is developed further in Can AI Search Visibility Predict Revenue Growth? How We Plan to Test the Relationship and How Investors Could Backtest AI Search Signals Against Revenue, Analyst Estimates and Stock Performance.
What This Article Does Not Claim
This article does not claim that:
- AI recommendation visibility predicts revenue;
- AI recommendation visibility predicts earnings;
- AI recommendation visibility predicts analyst revisions;
- AI recommendation visibility predicts stock returns;
- the initial 25-company signal is statistically validated as an investment factor;
- positive candidates are undervalued;
- negative candidates are overvalued;
- any listed company should be bought, sold, held, shorted, or avoided.
The purpose is narrower: define in advance what evidence would count against the hypothesis.
Want the full Authority Index
The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.
Methodology Context
The current V0 investor-signal methodology is based on the preserved LLM Authority Index research corpus containing:
- 68,203 raw observations;
- 310,114 raw citation-array entries;
- 624,698 company-array entries, including explicit not-mentioned company records;
- 1,278 explicit extraction-failed observations, excluded from valid AI visibility denominators;
- 5,976 exact cross-vertical duplicate observations identified in QA;
- 406 company/entity signals in the current momentum table;
- 25 mapped public parents in the first investor panel.
The V0 signal compares matched normalized prompts on the same AI platform family between a base month and September 2026. The primary measure is change in recommendation coverage. The current public-company watch rule combines magnitude, prompt-clustered exploratory interval direction, and platform breadth.
The full construction rules are documented in How We Measure AI Commercial Momentum.
Limitations of This Falsification Framework
This framework itself will require refinement as more data become available.
Current limitations include:
- only about three months of core longitudinal AI data are available;
- the first public panel contains only 25 mapped public parents;
- the panel is concentrated in consumer-facing financial, insurance, lending, healthcare, brokerage, and related categories;
- exact downstream outcome datasets have not yet been integrated for every company;
- segment exposure weights are not yet fully modeled;
- a formal threshold for the minimum economically meaningful incremental prediction gain has not yet been locked;
- market-return validation will require careful sector, factor, and risk adjustment;
- model and platform updates may change AI behavior over time.
Those limitations should make the project more conservative, not less falsifiable.
As the dataset grows, future articles should specify tighter minimum effect sizes, sample-size requirements, and holdout designs before evaluating later outcomes.
Why a Failed Hypothesis Would Still Be Useful
A negative result would answer an important question.
AI search measurement is rapidly becoming a commercial discipline. Companies can already measure whether they are cited, mentioned, recommended, and ranked across AI systems. But measurement does not automatically imply investor relevance.
If longitudinal evidence eventually shows that recommendation momentum has little or no incremental relationship with downstream financial outcomes, investors should know that.
It would imply that AI recommendation visibility is primarily a marketing, brand, search, or channel-performance metric rather than a distinct financial leading indicator.
That conclusion would prevent overinterpretation of AI visibility scores and could be more valuable than an attractive but unsupported valuation model.
Conversely, if a stable relationship does emerge, the fact that the failure conditions were published in advance will make that evidence stronger.
Related LLM Authority Index Research
- Can AI Search Signal Future Revenue Growth? The AI Commercial Momentum Hypothesis
- Can AI Recommendation Momentum Identify Emerging Company Trends? Initial Findings From 25 Public Companies
- AI Search as Alternative Data: Could AI Recommendations Become a Leading Indicator for Investors?
- How We Measure AI Commercial Momentum: Methodology for AI Investor Signals
- AI Investor Signal Tracker: Public Company AI Recommendation Momentum
- Can AI Search Visibility Predict Revenue Growth? How We Plan to Test the Relationship
- AI Visibility Market Divergence: Can AI Recommendation Momentum Reveal Information Not Yet Reflected in Investor Expectations?
- AI Recommendation Share vs. Market Share: Could the Gap Reveal Emerging Commercial Strength or Weakness?
- Does Cross-Platform AI Visibility Matter? Measuring Recommendation Portability and Platform Concentration Risk
- How Investors Could Backtest AI Search Signals Against Revenue, Analyst Estimates and Stock Performance
- The Persistence-Portability Gap
External References
- Michael Iannelli and Alan Ai. From Prompt to Purchase: How AI Brand Recommendations Move Consumers on the Open Web. 2026. https://arxiv.org/abs/2606.10907
- Maximilian Kaiser and Christian Schulze. Frontiers: ChatGPT Referrals to E-Commerce Websites: How Do LLMs Compare Against Traditional Channels? Marketing Science, 2026. https://pubsonline.informs.org/doi/10.1287/mksc.2025.0489
- AIVO Editorial Board. Correcting LLM Equity Valuation: why we withdrew our Grüns figures and what replaces them. September 23, 2026. https://www.aivojournal.org/correcting-llm-equity-valuation-why-we-withdrew-our-gruns-figures-and-what-replaces-them/
Research Status
This hypothesis remains unvalidated.
The next step is not to strengthen the language. The next step is to collect future outcomes and test whether the frozen AI signals actually anticipated anything economically important.
If they did, the historical publication record will show it.
If they did not, the same record should make that equally clear.
Want the full Authority Index
The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.
Keep reading
Related articles
AI Investor Signals
Can AI Search Signal Future Revenue Growth? The AI Commercial Momentum Hypothesis
Read this blog on LLM Authority Index.
Read articleAI Investor Signals
How We Measure AI Commercial Momentum: Methodology for AI Investor Signals
Read this blog on LLM Authority Index.
Read articleAI Investor Signals
What Is AI Recommendation Momentum? Measuring How AI Recommendations Change Over Time
Read this blog on LLM Authority Index.
Read articleSee how the framework applies to your market.
Get an AI Visibility Market Intelligence Report and see how AI is shaping consideration, comparison, and recommendation in your category.