用多个角色AI模拟人类评审,让自动评估更贴近真实多维度判断。
Multi-Agent-as-Judge: Aligning LLM-Agent-Based Automated Evaluation with Multi-Dimensional Human Evaluation
- 构建多个具不同视角的AI评审员,从文献自动生成角色设定。
- 通过多智能体辩论生成多维度反馈,与专家评分高度一致。
- 适用于教育、医疗等复杂领域,提升自动化评估可信度。
绝大多数人类工作具有协作性,因此现实NLP应用的评估往往需要多维且反映多元人类视角。由于真人评审资源稀缺且成本高,新兴的LLM-as-a-judge范式为利用大模型代理模拟人类评审提供了可能。然而,现有方法存在两个局限:代理角色设定常随意设计,且框架难以泛化至其他任务。为此,我们提出MAJ-EVAL,一个可自动从相关文本(如研究论文)中构建多维度评审角色、实例化大模型代理并组织多智能体群组辩论以生成多维度反馈的评估框架。在教育与医疗领域的实验表明,MAJ-EVAL生成的评估结果比传统自动化指标及现有LLM-as-a-judge方法更接近人类专家评分。
原文摘要 · Abstract (English)
Nearly all human work is collaborative; thus, the evaluation of real-world NLP applications often requires multiple dimensions that align with diverse human perspectives. As real human evaluator resources are often scarce and costly, the emerging "LLM-as-a-judge" paradigm sheds light on a promising approach to leverage LLM agents to believably simulate human evaluators. Yet, to date, existing LLM-as-a-judge approaches face two limitations: persona descriptions of agents are often arbitrarily designed, and the frameworks are not generalizable to other tasks. To address these challenges, we propose MAJ-EVAL, a Multi-Agent-as-Judge evaluation framework that can automatically construct multiple evaluator personas with distinct dimensions from relevant text documents (e.g., research papers), instantiate LLM agents with the personas, and engage in-group debates with multi-agents to Generate multi-dimensional feedback. Our evaluation experiments in both the educational and medical domains demonstrate that MAJ-EVAL can generate evaluation results that better align with human experts' ratings compared with conventional automated evaluation metrics and existing LLM-as-a-judge methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。