arXiv:2409.10883cs.CL2024-09被引 5

无需参考文本,用对比打分评估会议摘要质量。

CREAM: Comparison-Based Reference-Free ELO-Ranked Automatic Evaluation for Meeting Summarization

  • 通过思维链与关键事实对齐,评估摘要的完整性和简洁性。
  • 采用ELO评分系统,实现不同模型或提示配置的客观比较。
  • 适合需要快速评估会议摘要生成模型的研究者使用。

大型语言模型(LLMs)推动了摘要自动评估方法的发展,提供了比人工评估更快、更经济的替代方案。然而,现有方法在处理长上下文摘要和基于对话的会议摘要等复杂任务时往往表现不佳。本文提出CREAM(基于对比的无参考ELO评分会议摘要自动评估框架),针对会议摘要评估的独特挑战,利用思维链推理与关键事实对齐,无需参考文本即可评估生成摘要的完整性与简洁性。通过ELO排名系统,该方法能有效比较不同模型或提示配置的质量,实现稳定可靠的自动化评估。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have spurred interest in automatic evaluation methods for summarization, offering a faster, more cost-effective alternative to human evaluation. However, existing methods often fall short when applied to complex tasks like long-context summarizations and dialogue-based meeting summarizations. In this paper, we introduce CREAM (Comparison-Based Reference-Free Elo-Ranked Automatic Evaluation for Meeting Summarization), a novel framework that addresses the unique challenges of evaluating meeting summaries. CREAM leverages a combination of chain-of-thought reasoning and key facts alignment to assess conciseness and completeness of model-generated summaries without requiring reference. By employing an ELO ranking system, our approach provides a robust mechanism for comparing the quality of different models or prompt configurations.

自动评估会议摘要无参考ELO评分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。