让大模型评委通过对比众包回答,提升判断全面性。
Crowd Comparative Reasoning: Unlocking Comprehensive Evaluations for LLM-as-a-Judge
- 引入众包响应与待评回答对比,挖掘深层细节
- 平均准确率提升6.7%,且生成更高质量推理链
- 适合需要高可靠评估的模型训练与评测场景
LLM-as-a-Judge通过生成思维链(CoT)进行自动评估,已广泛应用。然而,其可靠性受制于思维链难以捕捉全面深入细节,常导致评估不完整。现有方法主要依赖多数投票或扩展评判标准,无法有效解决此问题。本文提出基于众包的对比评估方法,引入额外众包响应与候选回答进行对比,从而揭示候选回答中更深层、更全面的信息,有效引导大模型生成更详尽的思维链判断。大量实验表明,该方法在五个基准上平均准确率提升6.7%。此外,生成的思维链质量更高,有利于裁判模型蒸馏,并在监督微调(SFT)中的拒绝采样环节表现更优,即“众包拒绝采样”,显著提升微调效率。分析证实,本文生成的思维链更具全面性和高质量,且评估准确率随推理规模扩大而提升。
原文摘要 · Abstract (English)
LLM-as-a-Judge, which generates chain-of-thought (CoT) judgments, has become a widely adopted auto-evaluation method. However, its reliability is compromised by the CoT reasoning's inability to capture comprehensive and deeper details, often leading to incomplete outcomes. Existing methods mainly rely on majority voting or criteria expansion, which is insufficient to address the limitation in CoT. We propose Crowd-based Comparative Evaluation, which introduces additional crowd responses to compare with the candidate responses, thereby exposing deeper and more comprehensive details within the candidate responses. This process effectively guides LLM-as-a-Judge to provide a more detailed CoT judgment. Extensive experiments demonstrate that our approach enhances evaluation reliability, achieving an average accuracy gain of 6.7% across five benchmarks. Moreover, our method produces higher-quality CoTs that facilitate judge distillation and exhibit superior performance in rejection sampling for supervised fine-tuning (SFT), referred to as crowd rejection sampling, thereby enabling more efficient SFT. Our analysis confirms that CoTs generated by ours are more comprehensive and of higher quality, and evaluation accuracy improves as inference scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。