统一评估图文摘要的文本质量、图文对齐和视觉多样性,更真实反映多模态总结效果。
Measuring What Matters Beyond Text: Evaluating Multimodal Summaries by Quality, Alignment, and Diversity

- 构建三维度评估框架:文本质量、图文相关性、图像集合多样性
- 通过人类偏好校准,发现事实一致性最关键,视觉多样性为补充信号
- 无需参考文本,适合对比不同多模态摘要生成模型
多模态大语言模型(MLLM)推动了多模态摘要生成(MSMO),即从多源异构数据中生成简洁文本并配以关键视觉内容。然而现有评估方法碎片化,文本质量、图文对齐和视觉多样性常被孤立使用单模态指标,难以反映模态协同的忠实度与实用性。为此,我们提出MM-Eval统一评估框架,整合三项核心:(1) 文本质量,采用OpenFActScore评估事实一致性,G-Eval评估连贯性、流畅性与相关性;(2) 图文相关性,基于MLLM-as-a-judge方法;(3) 图像集多样性,使用截断CLIP熵量化。通过mLLM-EVAL新闻基准数据集训练学习聚合模型,使各组件贡献与人类偏好对齐。分析表明该场景下以文本为主导,事实一致性是整体感知质量的关键决定因素,而图文相关性与视觉多样性提供互补信息。相比启发式聚合基线,MM-Eval表现更优,提供可解释、弱依赖参考的多模态摘要比较工具。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have facilitated Multimodal Summarization with Multimodal Output (MSMO), wherein systems generate concise textual summaries accompanied by salient visuals from multimodal sources. However, current MSMO evaluation remains fragmented: text quality, image-text alignment, and visual diversity are typically assessed in isolation using unimodal metrics, making it difficult to capture whether the modalities jointly support a faithful and useful summary. To address this gap, we introduce MM-Eval, a unified evaluation framework that integrates assessments of textual quality, cross-modal alignment, and visual diversity. MM-Eval comprises three components: (1) text quality, measured using OpenFActScore for factual consistency and G-Eval for coherence, fluency, and relevance; (2) image-text relevance, evaluated via an MLLM-as-a-judge approach; and (3) image-set diversity, quantified using Truncated CLIP Entropy. We calibrate MM-Eval through a learned aggregation model trained on the mLLM-EVAL news benchmark, aligning component contributions with human preferences. Our analysis reveals a text-dominant hierarchy in this setting, where factual consistency acts as a critical determinant of perceived overall quality, while visual relevance and diversity provide complementary signals. MM-Eval improves over heuristic aggregation baselines and provides an interpretable, reference-weak framework for comparative evaluation of multimodal summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。