The Consensus Index: What Happens When We Stop Asking One Reviewer Who Is Best?

Learn what a Consensus Index measures in AI recommendations, why agreement is not truth, and how citations, disagreement, and source overlap matter.

Measurement15 minutesUpdated Aug 24, 2026By Mark Huntley, J.D.

Answer Capsule

A Consensus Index measures how multiple AI systems independently answer the same commercial question, then compares their recommendations, rankings, disagreements and cited sources. Its purpose is not to declare AI consensus to be truth. It is to make machine opinion measurable. A rigorous Consensus Index should account for source overlap, model disagreement, citation provenance and changes over time rather than simply counting how many AI platforms chose the same winner.

For decades, product-review publishing has revolved around a surprisingly simple question:

Who does the reviewer think is best?

Sometimes the reviewer is an individual.

Sometimes it is a testing laboratory.

Sometimes it is an editorial team.

Sometimes it is an affiliate publisher with a sophisticated methodology page and a financial relationship with many of the companies being ranked.

Regardless of who performs the analysis, the final output is usually the same:

one publication, one methodology, one ranking.

Artificial intelligence gives us the ability to ask a different question.

Instead of asking:

“Which medical alert system does this reviewer think is best?”

we can ask:

“What do the major AI systems currently believe is best—and how much do they agree?”

That is the idea behind what I call a Consensus Index.

But the ranking itself is only the beginning.

The more interesting information is hidden underneath it.

What Is a Consensus Index?

A Consensus Index is a structured measurement of how multiple AI systems respond independently to the same defined question.

For example:

“What are the best medical alert systems for a senior living alone in the United States?”

Instead of asking one AI once, the methodology can query multiple systems under controlled conditions and compare:

  • which brands appear,
  • which brands do not,
  • their ranking positions,
  • the reasons given for each recommendation,
  • the cited evidence,
  • the sources repeatedly appearing across systems,
  • areas of disagreement,
  • and changes in those answers over time.

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

The resulting index does not answer:

“Which product is objectively the best?”

It answers a different and increasingly important question:

“What is the current machine consensus about this category?”

That distinction is essential.

Consensus Is Not Truth

If six out of seven AI systems recommend the same company, it is tempting to write:

“Six AIs independently determined that Company A is the best.”

That would be too strong.

Those systems are not seven completely independent human experts locked in separate rooms.

They may have been trained on overlapping public information.

Their retrieval systems may find many of the same highly visible pages.

Those pages may themselves rely on one another.

Company marketing claims may have propagated through multiple publishers before eventually appearing in AI-generated answers.

So if six models agree, what we have demonstrated is machine agreement.

We have not necessarily demonstrated six independent confirmations of reality.

That is why I believe any serious AI Consensus Index must include provenance analysis.

The question is not merely:

“How many systems agreed?”

It is also:

“How independently supported was that agreement?”

The Old Idea Behind a New Measurement

There is nothing new about the idea that aggregating multiple judgments can produce useful information.

In 1907, statistician Francis Galton famously analyzed hundreds of estimates of an ox's weight submitted at a livestock competition.

The median estimate of the crowd came remarkably close to the actual result.[1]

That observation eventually became part of what is broadly described as the wisdom of crowds.

But there is an important condition attached to crowd wisdom:

diversity and independence matter.

Research by Lu Hong and Scott Page demonstrated mathematically that groups containing diverse approaches to problem solving can outperform groups composed only of individually high-performing problem solvers under certain conditions.[2]

Machine-learning research has reached a related conclusion.

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

Ensemble methods often improve predictive performance by combining multiple models—but their usefulness depends heavily on the models not simply making the same errors.[3]

That principle maps surprisingly well onto AI recommendation analysis.

Seven identical systems repeating the same source are not particularly interesting.

Seven substantially different systems arriving at similar conclusions using a varied evidence base are much more informative.

This suggests that an AI Consensus Index should measure not merely agreement, but independent agreement.

Why AI Makes Consensus Newly Observable

Before generative AI, measuring the collective opinion of the information ecosystem was difficult.

You could examine:

  • Google rankings,
  • product reviews,
  • Reddit threads,
  • Amazon reviews,
  • YouTube videos,
  • consumer surveys,
  • social mentions,
  • expert testing,
  • and traditional media.

But converting all of that into one coherent representation of what the internet “believes” required enormous manual analysis.

Large language models effectively perform synthesis continuously.

When a user asks an AI system:

“What are the best medical alert systems for someone living alone?”

the system attempts to compress a large body of information into a useful answer.

Modern retrieval-augmented systems can supplement model knowledge with external documents during answer generation.

The foundational Retrieval-Augmented Generation research by Lewis and colleagues demonstrated the value of combining generated responses with externally retrieved information for knowledge-intensive tasks.[4]

Today, generative search products increasingly expose some of the sources behind their answers.

That creates something new for researchers and marketers:

an observable output of machine synthesis.

We can now compare those outputs systematically.

The Aging in Place Index Experiment

This is what we began testing with Aging in Place Index.

Rather than commissioning another traditional review of medical alert systems, we asked multiple AI platforms to evaluate the market and then documented where their conclusions converged and diverged.

Our analysis of medical alert systems for seniors living alone produced an overall consensus ranking.

But the ranking was not the most interesting part.

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

The interesting questions were:

  • Why did Medical Guardian repeatedly appear near the top?
  • Which models preferred Bay Alarm Medical?
  • Where did LifeFone gain or lose support?
  • Which models surfaced weaknesses ignored by others?
  • Which publications appeared repeatedly in the supporting citations?
  • Were those publications genuinely independent?
  • Were first-party corporate sources dominating the evidence base?
  • Were certain facts repeated widely despite originating from only one source?

Those questions transform a “best products” article into something closer to an information-system audit.

We are no longer merely reviewing the products.

We are reviewing the information environment from which recommendations emerge.

Why Disagreement May Be More Valuable Than Agreement

Most ranking publications hide uncertainty.

An editor eventually has to choose:

  1. Company A
  2. Company B
  3. Company C

That presentation creates the appearance of certainty even when the underlying evidence is messy.

A Consensus Index can do the opposite.

It can explicitly expose disagreement.

For example:

Four systems rank Brand A first. Two rank Brand B first. One does not recommend Brand A at all.

That is valuable information.

Now ask why.

Perhaps one model places greater weight on price.

Another prioritizes response times.

Another finds a recent independent test.

Another relies heavily on manufacturer information.

Another surfaces a consumer complaint pattern.

The disagreement itself identifies which variables are unstable inside the information ecosystem.

For a consumer, that produces a more honest review.

For a company, it produces competitive intelligence.

For an LLM optimization team, it identifies where the brand narrative is fragmented.

The Truthfulness Problem Makes Transparency Essential

Large language models are extremely useful synthesis systems.

They are not infallible arbiters of fact.

Research has repeatedly demonstrated that language models can produce confident but inaccurate responses.

The TruthfulQA benchmark developed by Stephanie Lin, Jacob Hilton and Owain Evans specifically tested whether language models generate truthful answers to questions where humans commonly hold misconceptions.[5]

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

Their research showed that larger models do not automatically eliminate the problem of reproducing false or misleading beliefs present in their information environment.

A broader academic literature now examines what is commonly called hallucination in natural-language generation.[6]

This has an important implication for consensus analysis:

If several AI systems repeat the same falsehood, agreement does not transform it into truth.

This is why the underlying citations matter.

A transparent index should allow the reader to move from:

Conclusion

to:

Model

to:

Claim

to:

Citation

to:

Original source

wherever possible.

The more of that chain we can expose, the more useful the index becomes.

The Source Concentration Problem

Consider this hypothetical result:

Brand A

Recommended by:

  • ChatGPT
  • Gemini
  • Claude
  • Perplexity
  • Google AI Mode
  • Copilot

At first glance:

6/6 consensus.

Very compelling.

Now inspect the evidence.

Suppose the citation network looks like this:

AI System 1 → Review Site A

AI System 2 → Review Site A

AI System 3 → Review Site B

Review Site B → Review Site A

AI System 4 → Review Site C

Review Site C → Manufacturer

AI System 5 → Manufacturer

AI System 6 → Review Site A

The apparent six-way agreement may actually rest on two primary information sources.

That should affect our confidence in the result.

This is one reason I believe future Consensus Index methodology should incorporate a Source Concentration measure.

A recommendation backed by:

  • government data,
  • independent testing,
  • customer research,
  • multiple unrelated publishers,
  • technical specifications,
  • and professional evaluation

should be interpreted differently from a recommendation repeated across six websites that all originate from the same press release.

The models may agree equally in both examples.

The underlying evidence diversity is radically different.

A Better Way to Measure AI Consensus

A mature Consensus Index could eventually contain several separate metrics.

1. Model Consensus

How many AI systems recommend the entity?

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

Example:

6 of 7 systems

2. Average Recommendation Position

Where does the company appear when included?

Example:

Average rank: 1.8

3. Rank Variance

How consistently is the company ranked?

A brand that ranks:

1, 1, 2, 1, 2

has a different consensus profile from one ranked:

1, 7, 2, 6, 1

even if the average position is similar.

4. Citation Diversity

How many genuinely distinct sources support the recommendation?

5. Source Concentration

How much of the evidence ultimately depends upon a relatively small number of publications?

6. First-Party Dependency

What percentage of the supporting evidence originates from the company being evaluated?

A product whose recommendations depend primarily on manufacturer claims deserves a different confidence level from one supported by independent evidence.

7. Contradiction Rate

How frequently do credible sources disagree about material facts?

For example:

  • battery life,
  • pricing,
  • response time,
  • cancellation requirements,
  • range,
  • equipment fees.

8. Consensus Persistence

Does the ranking remain stable over multiple measurement periods?

One month of agreement is interesting.

Twelve months of agreement is something else entirely.

The Longitudinal Dataset May Become More Valuable Than the Ranking

This is where I believe the Consensus Index concept becomes particularly interesting.

Anyone can run seven prompts today.

Very few organizations will possess a clean dataset showing how machine recommendations changed every month for three years.

Imagine tracking a category from January 2027 through January 2030.

You could observe:

  • when a new company first entered machine recommendations,
  • when an established leader began declining,
  • which source first introduced a claim later repeated by other systems,
  • how quickly new product information propagated,
  • whether negative reporting altered machine sentiment,
  • whether a brand's traditional visibility preceded or followed AI visibility,
  • and whether the citation graph became more concentrated or more diverse.

The historical dataset cannot easily be reconstructed later.

Once January 2027 is gone, you cannot necessarily ask a model in 2030:

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

“Exactly what would you have recommended three years ago under the retrieval environment that existed at the time?”

The information environment will have changed.

That makes repeated measurement a potential data moat.

From Consensus to Citation Centrality

Once we preserve citations across hundreds or thousands of AI queries, another pattern should emerge.

Some sources will appear far more often—and in more consequential positions—than others.

That brings us to the next important measurement concept:

Citation Centrality.

Citation Centrality asks:

Which sources occupy the most influential positions inside the evidence networks shaping AI recommendations?

A website may have enormous traditional domain authority while rarely influencing a particular class of AI recommendations.

Another page may be relatively unknown to traditional marketers but repeatedly appear behind machine conclusions in one valuable niche.

Raw citation count will not be enough.

We will need to examine:

  • cross-model appearance,
  • proximity to recommendations,
  • persistence,
  • source independence,
  • category specificity,
  • and position within citation networks.

Read: Citation Centrality — Finding the Sources That Actually Shape AI Recommendations

And From Citation Centrality to Brand Centrality

There is a corresponding question for companies.

Once we identify the most influential sources in a category, we can ask:

Which brands are most strongly represented across those high-centrality sources?

I call this Brand Centrality.

Consider two competitors.

Company A

Appears on 500 websites.

Company B

Appears on 150 websites.

Traditional mention analysis might conclude that Company A dominates.

But suppose Company B appears positively on eight of the ten sources most consistently associated with AI recommendations.

Company A appears on only two.

Company B may possess substantially greater Brand Centrality even while having fewer total mentions.

That might explain why AI systems recommend Company B more often.

Read: Brand Centrality — Measuring a Company's Position Inside the AI Evidence Graph

What This Means for Product Reviews

I do not believe consensus publishing eliminates traditional product testing.

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

It shouldn't.

A laboratory can measure things an LLM cannot.

A human reviewer can experience:

  • comfort,
  • build quality,
  • usability,
  • customer service,
  • installation friction,
  • sound quality,
  • emotional response,
  • real-world reliability.

Those forms of firsthand experience remain valuable.

The Consensus Index adds another dimension.

Instead of asking a reviewer to replace the information ecosystem, we can ask:

What does the information ecosystem currently conclude?

And then inspect how that conclusion was formed.

The future review page may therefore contain multiple layers:

Human testing

What happened when someone actually used the product?

Independent evidence

What did laboratories, regulators, customers and professional evaluators find?

Machine consensus

What are the major AI systems currently recommending?

Provenance

What information is producing those recommendations?

Disagreement

Where do humans, sources and machines conflict?

Historical movement

How has that consensus changed?

That is a far richer information product than another generic “10 Best” article.

Why This Could Create Genuine Information Gain

There is an important difference between summarizing existing information and measuring existing information.

Suppose ten websites already describe a product.

An eleventh site rewrites those descriptions.

Little new information has been created.

But suppose the eleventh site asks seven generative systems the same standardized questions, records the outputs, normalizes them, calculates agreement, maps the citations and publishes the results.

The underlying product facts may not be new.

But the following facts are new:

“Six of seven systems recommend this company.”

“Its average recommendation position is 1.7.”

“Consensus increased from four platforms to six during the previous quarter.”

“Three publications account for 64% of its citation support.”

“Models agree on product quality but disagree materially on price.”

Those are new measurements of the information environment.

That is the information gain.

Generative Engine Optimization Research Supports Measuring Visibility

The academic field surrounding generative-engine visibility is still young.

One influential early paper, GEO: Generative Engine Optimization, examined whether modifications to source content could change how prominently that content appeared in generative-engine responses.[7]

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

The researchers showed that techniques involving citations, statistics, quotations and other content characteristics could materially affect source visibility.

That research does not demonstrate that our Consensus Index methodology is correct.

Nor does it establish that any particular AI recommendation is truthful.

What it does demonstrate is something foundational:

Visibility inside generative responses can be treated as a measurable phenomenon.

Once something is measurable, it can be compared.

Once it can be compared, it can be tracked.

Once it can be tracked longitudinally, patterns become visible.

That is the conceptual foundation of LLM Authority Index.

Consensus Should Be Reproducible

If a publication calls something an index, readers should be able to understand how the result was produced.

That means methodological disclosure matters.

Where practical, a Consensus Index should publish:

  • the question being measured,
  • prompt wording,
  • date of collection,
  • models or systems queried,
  • geographic assumptions,
  • ranking normalization rules,
  • scoring methodology,
  • exclusions,
  • underlying citations,
  • known limitations,
  • correction policies,
  • and historical revisions.

There will always be practical constraints.

Commercial AI systems change.

Outputs can vary.

Search results differ by location and time.

Prompts themselves influence answers.

Retrieval systems are not deterministic.

That doesn't make measurement useless.

It means the methodology needs to acknowledge uncertainty.

Financial indexes change.

Polling methodologies contain margins of error.

Scientific measurements contain noise.

The appropriate response is not to abandon measurement.

It is to disclose how the measurement was made.

The Difference Between an Index and an Oracle

This may be the most important distinction in the entire framework.

An oracle claims:

“This is the answer.”

An index claims:

“This is what we measured.”

LLM Authority Index should be the second.

If five models recommend Company A, we report that five models recommend Company A.

If three models disagree, we report the disagreement.

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

If 70% of supporting evidence traces back to company-owned material, we disclose that.

If the recommendation changes next month, the historical record should show the change.

The objective is not to make machines look omniscient.

It is to make machine opinion inspectable.

Why Brands Should Care

Consumers are increasingly using generative systems during product discovery and decision-making.

That means companies face an entirely new reputation question:

What do the machines believe about us?

But simply monitoring the final answer isn't enough.

If a model says:

“Competitor A is better for enterprise customers,”

the company needs to know why.

Which sources support that belief?

How frequently does the claim appear?

Do multiple systems agree?

Does the recommendation persist?

Did one publication introduce the idea?

Is the conclusion supported independently?

Is outdated information responsible?

Is a competitor simply better represented across influential third-party sources?

Those questions turn AI monitoring into citation intelligence.

And citation intelligence is where LLM optimization starts becoming much more strategically useful.

How the Consensus Index Fits the Larger Framework

The Consensus Index is one component of a larger model we are developing at LLM Authority Index.

The Algorithmic Reciprocity Loop

Describes how machine recognition can stimulate human citation, which can create additional web authority and future machine discoverability.

Read: The Algorithmic Reciprocity Loop

Consensus Index

Measures what AI systems currently conclude and how strongly they agree.

Search Authority vs. Machine Authority

Examines why traditional search prominence and machine recommendation influence are related but not necessarily equivalent.

Read: Authority Is Splitting in Two

Citation Rating

Identifies the sources that disproportionately influence machine conclusions.

Read: Citation Rating

Brand Rating

Measures how strongly brands are positioned within the influential evidence network.

Read: Brand Rating

Machine Relations

Turns that intelligence into a communications discipline: understanding which external information environments matter and earning legitimate inclusion within them.

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

Read on CiteWorks Studios: Machine Relations — Why PR Is Becoming Citation Engineering

A New Kind of Review Publication

Traditional reviews ask:

What should I buy?

Consensus analysis adds:

What do the machines recommend?

Citation analysis asks:

Why do they recommend it?

Citation Centrality asks:

Which sources matter most to that conclusion?

Brand Centrality asks:

Which companies occupy the strongest positions within those sources?

Longitudinal measurement asks:

How is all of this changing?

Those questions collectively create something more valuable than an AI-written product review.

They create an observable map of machine belief.

That is what interests me.

Not replacing human judgment.

Not declaring AI consensus to be objective truth.

Not generating thousands of automated “best products” pages.

But measuring an information system that is becoming increasingly influential over what people discover, compare and ultimately choose.

The internet spent decades building tools for measuring Google's results.

We are only beginning to build equivalent tools for understanding how machines synthesize the web itself.

The Consensus Index is one place to start.

Sources and Research

1. Galton, Francis — “Vox Populi.” Nature, 1907.
https://doi.org/10.1038/075450a0

Galton's analysis is an early and influential example of aggregate estimates producing a surprisingly accurate collective result. It provides historical context for the idea of consensus measurement, although AI systems should not be treated as statistically independent human judges.

2. Hong, Lu & Page, Scott E. — “Groups of Diverse Problem Solvers Can Outperform Groups of High-Ability Problem Solvers.” Proceedings of the National Academy of Sciences, 2004.
https://doi.org/10.1073/pnas.0403723101

The study demonstrates the potential value of diversity in collective problem solving under defined conditions. It supports the broader principle that diversity matters when aggregating judgments.

3. Dietterich, Thomas G. — “Ensemble Methods in Machine Learning.” Multiple Classifier Systems, 2000.

A foundational overview of ensemble methods in machine learning. Ensemble performance depends not merely on combining multiple models but on obtaining useful diversity among their errors and predictions.

4. Lewis, Patrick et al. — “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.” NeurIPS, 2020.
https://arxiv.org/abs/2005.11401

Foundational research on combining language generation with externally retrieved documents. It provides technical context for why source retrieval is important when examining modern generative information systems.

5. Lin, Stephanie; Hilton, Jacob; Evans, Owain — “TruthfulQA: Measuring How Models Mimic Human Falsehoods.” 2021/2022.
https://arxiv.org/abs/2109.07958

TruthfulQA demonstrates that language models can reproduce common misconceptions and falsehoods. It supports the methodological principle that model agreement should never automatically be treated as factual verification.

6. Ji, Ziwei et al. — “Survey of Hallucination in Natural Language Generation.” ACM Computing Surveys, 2023.
https://arxiv.org/abs/2202.03629

A broad review of hallucination research in natural-language generation, useful for understanding why transparency, external sourcing and verification remain essential.

7. Aggarwal, Pranjal et al. — “GEO: Generative Engine Optimization.” arXiv:2311.09735; presented at KDD 2024.
https://arxiv.org/abs/2311.09735

One of the foundational studies investigating source visibility inside generative-engine responses. The research provides support for treating generative visibility as something that can be systematically measured and influenced.

8. Google Search Central — “AI Features and Your Website.”
https://developers.google.com/search/docs/appearance/ai-features

Google's publisher documentation provides primary-source context for how web content can become eligible to appear within Google's AI-powered search features.

Methodology and Disclosure

The Consensus Index is a measurement framework being developed by Mark Huntley and LLM Authority Index.

It should not be interpreted as evidence that agreement among AI systems establishes objective truth.

AI systems may share:

  • training information,
  • retrieval sources,
  • underlying factual errors,
  • publisher biases,
  • commercial claims,
  • and broader web consensus.

For this reason, LLM Authority Index distinguishes between model agreement and evidence independence.

Our working hypothesis is that the most useful consensus measurement will combine recommendation frequency with source diversity, citation provenance, contradiction analysis and repeated longitudinal observation.

The methodology will continue to evolve as we collect additional data.

Want the full Authority Index

The paid deep-dive adds competitor threat profiles, the gap matrix, citation failure map, platform-by-platform recovery roadmap, and client-specific economic modeling.

See how the framework applies to your market.

Get an AI Market Intelligence Report and see how AI is shaping consideration, comparison, and recommendation in your category.