大规模评估大模型对话摘要能力,揭示推理与规模的影响
A Large-Scale Multi-Dimensional Empirical Study of LLMs for Conversation Summarization

- 构建涵盖6类场景的1800条对话基准,支持长达32k token输入
- 采用双向事实核验框架,量化评估摘要的完整、简洁与忠实性
- 评测28个模型,为实际部署提供选型依据
尽管大语言模型在对话摘要任务上取得显著进展,其评估仍受限于场景、输入长度和样本量不足,且现有基准常忽略前沿推理系统与高效小型模型,缺乏细粒度多维度分析。为此,我们提出OmniCSEval统一基准,包含1800条跨六类真实场景的多样化对话,上下文长度覆盖128至32,000个标记。为实现细粒度评估,采用双向事实核验框架,结合关键事实匹配以评估完整性和简洁性,以及摘要事实验证以评估忠实性。为确保评估可靠性,建立人-大模型协同的事实提取流程,并使用多大模型共识验证器进行摘要事实分解。基于此框架,我们在四个按推理能力与模型规模划分的类别中评估了28个大模型。广泛实证研究揭示了当前大模型在跨场景任务中的核心挑战,以及推理能力与模型规模的影响,同时分析了推理模型的效率与适应性,为实际部署提供指导。
原文摘要 · Abstract (English)
Despite the significant advancement of LLMs in conversation summarization, their evaluation remains limited by insufficient scenarios, input lengths, and sample sizes. Furthermore, existing benchmarks often omit frontier reasoning systems and efficient small models, or lack fine-grained, multi-dimensional assessments. To bridge these gaps, we propose OmniCSEval, a unified benchmark comprising 1,800 diverse conversations across six real-world scenarios, featuring context lengths ranging from 128 to 32k tokens. For fine-grained evaluation, we employ a bidirectional fact-checking framework that integrates key fact matching to assess completeness and conciseness, alongside summary fact verification to evaluate faithfulness. To ensure reliable assessment, we establish a human-LLM collaborative pipeline for key fact extraction and a multi-LLM consensus verifier for summary fact decomposition. Leveraging this framework, we evaluate 28 LLMs across four distinct categories grouped by reasoning capability and model scale. Our extensive empirical study reveals critical insights regarding the cross-scenario challenges current LLMs continue to face, the impacts of reasoning and scale, and the efficiency and adaptability of reasoning models. We also provide guidance for system selection in real-world deployments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。