arXiv:2503.15272cs.CLcs.AI2025-03NAACL被引 12

多智能体协作提升生成内容的忠实度,通过迭代修正消除事实错误。

MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration

  • 用多个不同类型的LLM协作检测和批评生成内容中的错误。
  • 在三个摘要数据集和长文本问答上,忠实度显著提升。
  • 将批评与修正视为重排序任务,性能优于传统生成方式。

多智能体协作在推理任务中表现优异,但在摘要、问答等长文本生成任务中仍待探索。本文将多智能体多模型推理拓展至生成任务,聚焦于通过修正提升生成内容的忠实度,即删除事实性错误。研究了多个实例和多种大型语言模型(LLMs)在错误检测、批判不忠实语句及基于批判进行修正等子任务中的协同效果。设计了各子任务的内在评估指标,结果表明多智能体(多个实例)与多模型(多样化LLM类型)均有助于提升错误检测与批判能力。此外,将批判与修正重构为重排序任务,而非生成任务,可进一步提升多智能体性能。基于上述发现,提出最终方案MAMM-Refine:多智能体多模型协同修正。该方法在三个摘要数据集及长文本问答任务上均取得显著提升,验证了其有效性与通用性。

原文摘要 · Abstract (English)

Multi-agent collaboration among models has shown promise in reasoning tasks but is underexplored in long-form generation tasks like summarization and question-answering. We extend multi-agent multi-model reasoning to generation, specifically to improving faithfulness through refinement, i.e., revising model-generated outputs to remove factual inconsistencies. We investigate how iterative collaboration among multiple instances and types of large language models (LLMs) enhances subtasks in the refinement process, such as error detection, critiquing unfaithful sentences, and making corrections based on critiques. We design intrinsic evaluations for each subtask, with our findings indicating that both multi-agent (multiple instances) and multi-model (diverse LLM types) approaches benefit error detection and critiquing. Additionally, reframing critiquing and refinement as reranking rather than generation tasks improves multi-agent performance. We consolidate these insights into a final "recipe" called Multi-Agent Multi-Model Refinement (MAMM-Refine), where multi-agent and multi-model collaboration significantly boosts performance on three summarization datasets as well as on long-form question answering, demonstrating the effectiveness and generalizability of our recipe.

多智能体生成忠实度大模型协作摘要优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。