arXiv:2602.08149cs.CLcs.AI2026-02

为对话摘要错误设计分层评估框架,提升自动评价精度。

DIAL-SUMMER: A Structured Evaluation Framework of Hierarchical Errors in Dialogue Summaries

  • 构建两级错误分类体系:对话级与句内转述级
  • 发现中段对话最易遗漏,结尾常出现虚构信息
  • 适合评估对话摘要质量或改进大模型评判能力的研究者

对话是人类交流的主要形式,自动生成对话摘要极具价值(如会议要点回顾、客服与用户对话复盘)。现有对话摘要评估方法忽视了任务特有复杂性:一是多说话人分散讨论内容向摘要句子的结构转换,二是第一/二人称叙述转为摘要中的第三人称标准化表达。本文提出DIAL-SUMMER框架,建立涵盖对话级(关注整体发言者/轮次)与句内转述级(关注单轮信息)的错误分类体系,并构建人工标注的语料库。实证分析显示:对话中段轮次最易遗漏,外部幻觉多出现在摘要末尾。通过测试LLM判别能力,验证了数据集挑战性与分类体系稳健性,凸显未来需提升大模型在该任务上的表现。代码与推理数据集即将发布。

原文摘要 · Abstract (English)

Dialogues are a predominant mode of communication for humans, and it is immensely helpful to have automatically generated summaries of them (e.g., to revise key points discussed in a meeting, to review conversations between customer agents and product users). Prior works on dialogue summary evaluation largely ignore the complexities specific to this task: (i) shift in structure, from multiple speakers discussing information in a scattered fashion across several turns, to a summary's sentences, and (ii) shift in narration viewpoint, from speakers' first/second-person narration, standardized third-person narration in the summary. In this work, we introduce our framework DIALSUMMER to address the above. We propose DIAL-SUMMER's taxonomy of errors to comprehensively evaluate dialogue summaries at two hierarchical levels: DIALOGUE-LEVEL that focuses on the broader speakers/turns, and WITHIN-TURN-LEVEL that focuses on the information talked about inside a turn. We then present DIAL-SUMMER's dataset composed of dialogue summaries manually annotated with our taxonomy's fine-grained errors. We conduct empirical analyses of these annotated errors, and observe interesting trends (e.g., turns occurring in middle of the dialogue are the most frequently missed in the summary, extrinsic hallucinations largely occur at the end of the summary). We also conduct experiments on LLM-Judges' capability at detecting these errors, through which we demonstrate the challenging nature of our dataset, the robustness of our taxonomy, and the need for future work in this field to enhance LLMs' performance in the same. Code and inference dataset coming soon.

对话摘要评估框架错误分析大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。