arXiv:2409.19898cs.CLcs.AI2024-09EMNLP被引 24

构建细粒度多维度摘要评估基准,提升大模型总结质量评测的全面性与客观性。

UniSumEval: Towards Unified, Fine-Grained, Multi-Dimensional Summarization Evaluation for LLMs

  • 设计覆盖多种文本领域和长度的评估场景,支持细粒度多维标注
  • 在九个主流大模型上测试,揭示其在不同情境下的表现差异
  • 结合AI辅助生成与人工标注,降低标注难度并提升一致性

现有摘要质量评估基准普遍存在输入场景单一、评估维度狭窄(如仅关注忠实性)以及主观性强、标注粗略等问题。为此,我们构建了UniSumEval评估基准,扩展了输入上下文范围(如领域、长度),并提供细粒度、多维度的标注。通过引入AI辅助数据生成,识别潜在幻觉文本,并帮助人工标注者降低细粒度标注任务难度。基于UniSumEval,我们对九个最新的语言模型进行了评估,揭示其在不同输入情境与评估维度下的表现差异。此外,还系统比较了当前最优的自动化摘要评估方法。相关数据集将公开于 https://github.com/DISL-Lab/UniSumEval-v1.0。

原文摘要 · Abstract (English)

Existing benchmarks for summarization quality evaluation often lack diverse input scenarios, focus on narrowly defined dimensions (e.g., faithfulness), and struggle with subjective and coarse-grained annotation schemes. To address these shortcomings, we create UniSumEval benchmark, which extends the range of input context (e.g., domain, length) and provides fine-grained, multi-dimensional annotations. We use AI assistance in data creation, identifying potentially hallucinogenic input texts, and also helping human annotators reduce the difficulty of fine-grained annotation tasks. With UniSumEval, we benchmark nine latest language models as summarizers, offering insights into their performance across varying input contexts and evaluation dimensions. Furthermore, we conduct a thorough comparison of SOTA automated summary evaluators. Our benchmark data will be available at https://github.com/DISL-Lab/UniSumEval-v1.0.

摘要评估大模型评测细粒度标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。