用多智能体协作提升大模型评分一致性与公正性
CollabEval: Enhancing LLM-as-a-Judge via Multi-Agent Collaboration
- 多个大模型协同评估,通过三阶段合作达成共识
- 在多个评测维度上显著优于单模型评分,抗干扰能力强
- 适合需要高可信度自动评分的场景,如内容审核、教育评测
大型语言模型(LLMs)已革新AI生成内容的评估方式,其中LLM-as-a-Judge范式日益流行。然而,现有单模型评估方法存在判断不一致和预训练数据固有偏见等问题。为此,我们提出CollabEval,一种新型多智能体评估框架,采用三阶段协同评估流程:初始评估、多轮讨论和最终判决。与依赖竞争辩论或单一模型评估的方法不同,CollabEval强调多个智能体间的协作,并通过策略性共识检查提升效率。大量实验表明,CollabEval在多个维度上持续优于单模型方法,即使个别模型表现不佳也保持稳健性能。该框架支持多种评估标准,同时通过协作设计保障高效性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have revolutionized AI-generated content evaluation, with the LLM-as-a-Judge paradigm becoming increasingly popular. However, current single-LLM evaluation approaches face significant challenges, including inconsistent judgments and inherent biases from pre-training data. To address these limitations, we propose CollabEval, a novel multi-agent evaluation framework that implements a three-phase Collaborative Evaluation process: initial evaluation, multi-round discussion, and final judgment. Unlike existing approaches that rely on competitive debate or single-model evaluation, CollabEval emphasizes collaboration among multiple agents with strategic consensus checking for efficiency. Our extensive experiments demonstrate that CollabEval consistently outperforms single-LLM approaches across multiple dimensions while maintaining robust performance even when individual models struggle. The framework provides comprehensive support for various evaluation criteria while ensuring efficiency through its collaborative design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。