arXiv:2502.00322cs.CLcs.IR2025-02NAACL被引 6

让多个AI角色各说各话,再由主控者整合出公正的辩论摘要。

MODS: Moderating a Mixture of Document Speakers to Summarize Debatable Queries in Document Collections

  • 用多个专用AI模拟不同文档立场,主控者按主题调度发言。
  • 在两个数据集上,关键话题覆盖率达90%以上,平衡性提升59%。
  • 适合需要客观综述争议性问题的研究者与决策者使用。

查询聚焦摘要(QFS)旨在生成回答查询的文档摘要。以往研究假设查询只有一个答案,忽略了具有争议性的问题(如‘法学院是否值得读?’)。本文提出争议性查询聚焦摘要(DQFS),任务是通过包含对立观点的文档生成全面且平衡的摘要,不偏袒任何一方。现有大模型方法存在两大缺陷:1)缺乏结构化内容规划,无法引导模型写出均衡摘要;2)对所有文档使用相同查询检索上下文,导致无法捕捉每篇文档的独特视角。为此,本文设计MODS——一种模仿人类研讨会议程的多大模型框架。将每个文档视为独立的‘发言人’(Speaker LLM),由一个‘主持人’(Moderator LLM)根据预设主题选择发言人,并为其定制查询以检索对应内容,再将各方观点汇总至详细大纲,形成内容计划,指导最终摘要生成。在ConflictingQA和新构建的DebateQFS数据集(来自Debatepedia的争议性问题)上的实验表明,MODS相比当前最优方法,在话题段落覆盖率和平衡性上提升38%-59%,基于新提出的引用指标。用户评估也证实其摘要更易读、更公正。

原文摘要 · Abstract (English)

Query-focused summarization (QFS) gives a summary of documents to answer a query. Past QFS work assumes queries have one answer, ignoring debatable ones (Is law school worth it?). We introduce Debatable QFS (DQFS), a task to create summaries that answer debatable queries via documents with opposing perspectives; summaries must comprehensively cover all sources and balance perspectives, favoring no side. These goals elude LLM QFS systems, which: 1) lack structured content plans, failing to guide LLMs to write balanced summaries, and 2) use the same query to retrieve contexts across documents, failing to cover all perspectives specific to each document's content. To overcome this, we design MODS, a multi-LLM framework mirroring human panel discussions. MODS treats documents as individual Speaker LLMs and has a Moderator LLM that picks speakers to respond to tailored queries for planned topics. Speakers use tailored queries to retrieve relevant contexts from their documents and supply perspectives, which are tracked in a rich outline, yielding a content plan to guide the final summary. Experiments on ConflictingQA with controversial web queries and DebateQFS, our new dataset of debate queries from Debatepedia, show MODS beats SOTA by 38-59% in topic paragraph coverage and balance, based on new citation metrics. Users also find MODS's summaries to be readable and more balanced.

摘要生成争议性问题多智能体大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。