arXiv:2607.21632cs.CL2026-07

用大模型互评方式,判断谁的回答更优。

A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models

论文配图:A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models
图 1 · 摘自论文原文
  • 让多个大模型匿名打分,比谁的答案更受认可。
  • 五种任务中,特定模型常被其他模型优先选择。
  • 适合评估多正确答案场景下的回答质量差异。

传统大模型评测依赖静态数据集和客观评分,难以区分多个合理回答在清晰度、完整性和实用性上的差异。本文提出一种基于共识的相对偏好评估框架,不以固定标准为基准,而是让一组多样化的大型语言模型对匿名生成的多个回答进行独立排序。通过结构化投票,将各模型的偏好聚合为相对智能指数(RII),反映某模型回答被其他模型青睐的频率。我们在编程、通用知识、安全性、逻辑推理和数学等多领域,使用五个顶尖大模型开展控制实验。结果显示跨领域存在稳定的偏好模式,部分模型在多数情况下更受同行青睐。但需强调,该结果体现的是模型间的偏好一致性,而非绝对正确性或人类判断。此框架为多有效解场景提供了可扩展、模型驱动的对比评估方法。虽不直接对应人类评价,已有研究表明模型共识可部分反映人类偏好,因而可作为代理信号。

原文摘要 · Abstract (English)

Traditional benchmarks for LLMs primarily rely on static datasets and objective scoring metrics, which often fail to capture differences in response quality when multiple answers are acceptable. In such settings, correctness alone is insufficient to distinguish between responses that vary in clarity, completeness, and usefulness. This paper introduces a consensus-based evaluation framework that measures relative preference among model-generated responses rather than absolute correctness. Instead of evaluating outputs against a fixed ground truth, we assess how a panel of diverse LLMs ranks anonymized candidate responses to the same prompt. This approach treats aggregate inter-model agreement as a proxy for perceived response quality under blind conditions. We conduct a controlled study using five state-of-the-art LLMs across multiple domains, including programming, general knowledge, safety, logical reasoning, and mathematics. Each model generates responses and independently ranks peer outputs through a structured voting process. Scores are aggregated into a Relative Intelligence Index (RII), representing how frequently a model's responses are preferred by other models. Our findings reveal consistent preference patterns across domains, with certain models more frequently ranked highly by their peers. However, we emphasize that these results reflect inter-model preference alignment rather than objective correctness or human judgment. This framework provides a scalable, model-driven method for comparative evaluation, offering an alternative perspective on response quality in scenarios where multiple valid answers exist. While not directly aligned with human evaluation, prior work suggests that aggregated model preferences can partially correlate with human judgments, motivating this as a proxy signal.

模型评估相对偏好共识机制LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。