对比推理与非推理模型,发现推理未必提升对话摘要质量。
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
- 对比推理与非推理大模型在三类对话摘要中的表现
- 推理模型常更啰嗦、事实错误更多,摘要更不简洁
- 适合关注对话摘要实际效果的研究者与应用开发者
对话摘要是一项具有重要实用价值的挑战性任务,广泛应用于客户服务、会议分析和对话式AI。尽管大语言模型在摘要任务中取得显著进展,但针对需要同时实现抽象与简洁性的对话场景,基于逐步推理架构(如OpenAI-o1和DeepSeek-R1)的长链思维(CoT)实现效果仍缺乏系统评估。本文首次对当前最先进的推理与非推理大模型,在通用、角色导向和查询导向三类对话摘要范式下进行全面系统评估。研究覆盖多种语言、领域和摘要长度,采用SAMSum、DialogSum、CSDS和QMSum等强基准数据集,结合基于大模型的自动指标与人类启发式评价标准。结果表明,与其它推理密集型任务趋势相反,显式分步推理并未持续提升对话摘要质量。相反,推理模型往往更冗长、存在事实矛盾,且摘要更不简洁,相比非推理模型表现更差。通过场景特异性分析与案例研究,进一步揭示了在复杂对话背景下,显式推理可能失效甚至产生负面影响的原因。本工作为当前推理大模型的局限性提供了新见解,并强调需针对真实场景设计针对性建模与评估策略。
原文摘要 · Abstract (English)
Dialogue summarization is a challenging task with significant practical value in customer service, meeting analysis, and conversational AI. Although large language models (LLMs) have achieved substantial progress in summarization tasks, the performance of step-by-step reasoning architectures-specifically Long Chain-of-Thought (CoT) implementations such as OpenAI-o1 and DeepSeek-R1-remains unexplored for dialogue scenarios requiring concurrent abstraction and conciseness. In this work, we present the first comprehensive and systematic evaluation of state-of-the-art reasoning LLMs and non-reasoning LLMs across three major paradigms-generic, role-oriented, and query-oriented dialogue summarization. Our study spans diverse languages, domains, and summary lengths, leveraging strong benchmarks (SAMSum, DialogSum, CSDS, and QMSum) and advanced evaluation protocols that include both LLM-based automatic metrics and human-inspired criteria. Contrary to trends in other reasoning-intensive tasks, our findings show that explicit stepwise reasoning does not consistently improve dialogue summarization quality. Instead, reasoning LLMs are often prone to verbosity, factual inconsistencies, and less concise summaries compared to their non-reasoning counterparts. Through scenario-specific analyses and detailed case studies, we further identify when and why explicit reasoning may fail to benefit-or even hinder-summarization in complex dialogue contexts. Our work provides new insights into the limitations of current reasoning LLMs and highlights the need for targeted modeling and evaluation strategies for real-world dialogue summarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。