用多角色辩论框架提升大模型评估的可靠性和可解释性
Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation
- 设计多智能体辩论机制,分角色协作评估大模型输出
- 支持并行辩护与预算控制迭代,显著提升评分一致性
- 适合关注评估公平性与成本效益的研究者和工程师
大型语言模型(LLM)的评估仍面临不一致、偏见和决策标准不透明的问题。本文提出一种成本感知的对抗式多智能体框架D3,通过角色专业化智能体(辩护人、裁判及可选陪审团)之间的结构化辩论,实现可靠且可解释的评估。D3包含两种互补协议:(1) 多辩护人单轮评估(MORE),每答案生成k个并行辩护以增强信号;(2) 单辩护人多轮评估(SAMRE),在显式令牌预算和收敛检查下迭代优化论点。我们构建了分数差距的概率模型,证明在温和假设下,回合间差距后验分布聚焦于真实差异,误排序概率趋近于零;且聚合k个辩护人能严格提升期望得分差距。实验覆盖MT-Bench、AlignBench和AUTO-J,结果显示与人工判断达到最先进的吻合度(准确率与Cohen's kappa),通过匿名化和角色多样化降低位置与冗长偏见,并在成本-精度权衡上表现优异。消融与定性分析验证了辩论、聚合与匿名性的贡献。D3为可信、可解释、成本敏感的LLM评估提供了系统性解决方案。
原文摘要 · Abstract (English)
The evaluation of Large Language Models (LLMs) remains challenging due to inconsistency, bias, and the absence of transparent decision criteria in automated judging. We present Debate, Deliberate, Decide (D3), a cost-aware, adversarial multi-agent framework that orchestrates structured debate among role-specialized agents (advocates, a judge, and an optional jury) to produce reliable and interpretable evaluations. D3 instantiates two complementary protocols: (1) Multi-Advocate One-Round Evaluation (MORE), which elicits k parallel defenses per answer to amplify signal via diverse advocacy, and (2) Single-Advocate Multi-Round Evaluation (SAMRE) with budgeted stopping, which iteratively refines arguments under an explicit token budget and convergence checks. We develop a probabilistic model of score gaps that (i) characterizes reliability and convergence under iterative debate and (ii) explains the separation gains from parallel advocacy. Under mild assumptions, the posterior distribution of the round-r gap concentrates around the true difference and the probability of mis-ranking vanishes; moreover, aggregating across k advocates provably increases expected score separation. We complement theory with a rigorous experimental suite across MT-Bench, AlignBench, and AUTO-J, showing state-of-the-art agreement with human judgments (accuracy and Cohen's kappa), reduced positional and verbosity biases via anonymization and role diversification, and a favorable cost-accuracy frontier enabled by budgeted stopping. Ablations and qualitative analyses isolate the contributions of debate, aggregation, and anonymity. Together, these results establish D3 as a principled, practical recipe for reliable, interpretable, and cost-aware LLM evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。