为有背景知识的讨论生成清晰摘要,解决引用不清、信息缺失问题。
What Are They Talking About? A Benchmark of Knowledge-Grounded Discussion Summarization
- 提出知识引导型讨论摘要新任务,分背景与观点两部分生成。
- 构建首个包含多粒度标注的基准数据集,覆盖新闻-讨论配对。
- 发现主流大模型仍难处理隐含引用和关键事实遗漏问题。
传统对话摘要仅关注对话内容,假设其本身已足够完整。但在基于共同背景的讨论中,参与者常省略上下文并使用隐含指代,导致不熟悉背景的读者难以理解摘要。为此,我们提出知识引导型讨论摘要(KGDS)任务,旨在生成补充背景摘要和澄清引用的观点摘要。为推动研究,我们构建了首个KGDS基准,包含新闻-讨论配对,并由专家创建多粒度真实标签以评估子摘要。我们还设计了一种分层评估框架,包含细粒度且可解释的指标。对12个先进大语言模型的广泛评估表明,KGDS仍是重大挑战:模型在背景摘要中常遗漏关键事实,保留无关信息;在观点摘要整合中难以解析隐含引用。
原文摘要 · Abstract (English)
Traditional dialogue summarization primarily focuses on dialogue content, assuming it comprises adequate information for a clear summary. However, this assumption often fails for discussions grounded in shared background, where participants frequently omit context and use implicit references. This results in summaries that are confusing to readers unfamiliar with the background. To address this, we introduce Knowledge-Grounded Discussion Summarization (KGDS), a novel task that produces a supplementary background summary for context and a clear opinion summary with clarified references. To facilitate research, we construct the first KGDS benchmark, featuring news-discussion pairs and expert-created multi-granularity gold annotations for evaluating sub-summaries. We also propose a novel hierarchical evaluation framework with fine-grained and interpretable metrics. Our extensive evaluation of 12 advanced large language models (LLMs) reveals that KGDS remains a significant challenge. The models frequently miss key facts and retain irrelevant ones in background summarization, and often fail to resolve implicit references in opinion summary integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。