用多智能体辩论法提升摘要忠实性评估的准确性
Faithful, Unfaithful or Ambiguous? Multi-Agent Debate with Initial Stance for Summary Evaluation
- 多个大模型智能体强制持不同立场辩论,激发更深入的错误发现
- 在真实数据集上识别出更多模糊情况,非模糊摘要评估更准
- 适合需要高精度摘要质检的研究者与工业应用
基于大语言模型的摘要忠实性评估方法常被文本流畅性误导,难以识别摘要错误。本文提出一种多智能体辩论机制:多个基于LLM的智能体被赋予初始立场(无论其真实信念如何),被迫为该立场寻找理由,通过多轮辩论达成共识。统一随机分配初始立场带来更大观点多样性,促进更有意义的辩论,从而识别更多错误。此外,通过对近期忠实性评估数据集的分析,我们发现摘要并不总是完全忠实或不忠实,因此引入“模糊性”新维度,并构建详细分类体系以识别此类特殊情形。实验表明,该方法能有效识别模糊情况,在非模糊摘要上表现更优。
原文摘要 · Abstract (English)
Faithfulness evaluators based on large language models (LLMs) are often fooled by the fluency of the text and struggle with identifying errors in the summaries. We propose an approach to summary faithfulness evaluation in which multiple LLM-based agents are assigned initial stances (regardless of what their belief might be) and forced to come up with a reason to justify the imposed belief, thus engaging in a multi-round debate to reach an agreement. The uniformly distributed initial assignments result in a greater diversity of stances leading to more meaningful debates and ultimately more errors identified. Furthermore, by analyzing the recent faithfulness evaluation datasets, we observe that naturally, it is not always the case for a summary to be either faithful to the source document or not. We therefore introduce a new dimension, ambiguity, and a detailed taxonomy to identify such special cases. Experiments demonstrate our approach can help identify ambiguities, and have even a stronger performance on non-ambiguous summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。