Paper proposes model-consensus framework to rank LLM responses without fixed ground truth
New evaluation method aggregates peer rankings to produce a Relative Intelligence Index, aiming to capture nuanced differences in response quality where multiple answers are valid.
1 source · cross-referenced
- A new arXiv paper introduces a consensus-based evaluation framework that ranks LLM responses by aggregating peer preferences rather than relying on fixed ground truth.
- The method uses a panel of five state-of-the-art LLMs to vote on anonymized responses across programming, general knowledge, safety, logical reasoning, and mathematics.
- Scores are compiled into a Relative Intelligence Index (RII), reflecting how often a model’s outputs are preferred by other models.
- Authors note the results reflect inter-model preference alignment, not objective correctness or human judgment.
A new paper on arXiv proposes a consensus-based framework for evaluating large language models by measuring relative preference among responses rather than absolute correctness. The approach replaces fixed ground truth with a voting process where a panel of diverse LLMs ranks anonymized candidate responses to identical prompts.
The study evaluates five state-of-the-art LLMs across five domains: programming, general knowledge, safety, logical reasoning, and mathematics. Each model generates responses and independently ranks peer outputs through a structured voting process.
Scores are aggregated into a Relative Intelligence Index (RII), which quantifies how frequently a model’s responses are preferred by other models. The authors emphasize that the RII reflects inter-model preference alignment rather than objective correctness or human judgment.
The framework is positioned as a scalable, model-driven alternative to traditional benchmarks, particularly in scenarios where multiple valid answers exist and correctness alone cannot distinguish nuanced differences in clarity, completeness, and usefulness.
While not directly aligned with human evaluation, the authors cite prior work suggesting that aggregated model preferences can partially correlate with human judgments, motivating the use of model consensus as a proxy signal.
- Jul 20, 2026 · Apple — Machine Learning Research
Apple releases LVSum, a benchmark for timestamp-aware long video summarization with 72 videos across 13 domains
Trust84 - Jul 14, 2026 · arXiv cs.CL
Researchers release CLIR-Bench to evaluate multimodal QA over irregular clinical time series
Trust79 - Jul 12, 2026 · Hamel Husain — applied AI engineering
Comparison finds automated evals correlate with human annotations in 100 traces
Trust79