arXiv:2509.15901cs.CLcs.AI2025-09EMNLP被引 4

用问答引导生成更准确、个性化的会议摘要,减少幻觉和遗漏。

Re-FRAME the Meeting Summarization SCOPE: Fact-Based Summarization and Personalization via Questions

  • 将摘要任务转为语义增强,通过提取关键事实并组织主题来构建摘要。
  • 在QMSum和FAME数据集上,幻觉和遗漏率降低2分(满分5分)。
  • 适合需要高准确性与个性化摘要的会议记录场景。

大语言模型进行会议摘要仍易出现幻觉、遗漏和无关内容。我们提出FRAME,一个模块化流程,将摘要重构为语义增强任务:提取并评分关键事实,按主题组织,并用于丰富摘要提纲。为实现个性化,引入SCOPE——一种‘边思考边回答’协议,要求模型先回答九个问题再选择内容。评估方面,提出P-MESA,一种多维度、无需参考文本的评估框架,可有效识别错误,与人工标注平衡准确率≥89%,与人类严重性评级相关系数r≥0.70。在QMSum和FAME数据集上,FRAME使幻觉与遗漏减少2/5分(以MESA衡量),SCOPE显著提升知识契合度与目标对齐性,优于仅用提示的方法。研究倡导重新思考摘要范式,以提升可控性、忠实性与个性化。

原文摘要 · Abstract (English)

Meeting summarization with large language models (LLMs) remains error-prone, often producing outputs with hallucinations, omissions, and irrelevancies. We present FRAME, a modular pipeline that reframes summarization as a semantic enrichment task. FRAME extracts and scores salient facts, organizes them thematically, and uses these to enrich an outline into an abstractive summary. To personalize summaries, we introduce SCOPE, a reason-out-loud protocol that has the model build a reasoning trace by answering nine questions before content selection. For evaluation, we propose P-MESA, a multi-dimensional, reference-free evaluation framework to assess if a summary fits a target reader. P-MESA reliably identifies error instances, achieving >= 89% balanced accuracy against human annotations and strongly aligns with human severity ratings (r >= 0.70). On QMSum and FAME, FRAME reduces hallucination and omission by 2 out of 5 points (measured with MESA), while SCOPE improves knowledge fit and goal alignment over prompt-only baselines. Our findings advocate for rethinking summarization to improve control, faithfulness, and personalization.

会议摘要大模型个性化幻觉抑制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。