首个多模态对话摘要元评估基准,解决自动评价缺乏标准的问题。
MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
- 构建跨模态关键信息过滤框架,提升数据质量
- 涵盖8个维度的人工评判,形成首个元评估基准
- 揭示现有评估方法对大模型摘要的误判问题
多模态对话摘要(MDS)是一项具有广泛应用的重要任务。为支持高效MDS模型的开发,可靠的自动评估方法可显著降低人力与成本开销。然而,这类方法依赖于基于人工标注的强元评估基准。本文提出MDSEval,首个针对MDS的元评估基准,包含图像分享对话、对应摘要及在八个明确质量维度上的人工判断。为确保数据质量与丰富性,我们设计了一种新颖的过滤框架,利用跨模态互斥关键信息(MEKI)。本工作首次系统识别并形式化了MDS特有的关键评估维度。我们对当前主流的自动评估方法进行了基准测试,发现其在区分先进多模态大模型生成的摘要方面存在局限,并易受多种偏差影响。
原文摘要 · Abstract (English)
Multimodal Dialogue Summarization (MDS) is a critical task with wide-ranging applications. To support the development of effective MDS models, robust automatic evaluation methods are essential for reducing both cost and human effort. However, such methods require a strong meta-evaluation benchmark grounded in human annotations. In this work, we introduce MDSEval, the first meta-evaluation benchmark for MDS, consisting image-sharing dialogues, corresponding summaries, and human judgments across eight well-defined quality aspects. To ensure data quality and richfulness, we propose a novel filtering framework leveraging Mutually Exclusive Key Information (MEKI) across modalities. Our work is the first to identify and formalize key evaluation dimensions specific to MDS. We benchmark state-of-the-art modal evaluation methods, revealing their limitations in distinguishing summaries from advanced MLLMs and their susceptibility to various bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。