无需参考摘要,用多模态问答评估视频摘要质量
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering

- 通过多模态问答直接对比候选摘要与原始视频
- 在800条摘要上验证,相关性优于现有方法
- 适合关注自动评估的视频摘要研究者
视频到文本摘要的评估方法仍不完善。传统基于n-gram重叠的指标和近期基于大语言模型的方法严重依赖人工编写的参考摘要,限制了实用性并难以捕捉细微语义。本文提出QEVA,一种无需参考的评估指标,通过多模态问答将候选摘要直接与源视频对比。QEVA从覆盖度、事实性和时序性三个维度评估摘要质量。我们还构建了新基准MLVU(VS)-Eval,源自MLVU数据集,包含200个视频生成的800条摘要,由当前先进视频-语言模型生成。该数据集建立了透明一致的评估框架。实验表明,相较于现有方法,QEVA在肯德尔等级相关系数τ_b、τ_c和斯皮尔曼等级相关系数ρ上均与人类判断具有更高相关性。我们希望本研究的基准与指标能推动视频到文本摘要研究的发展,并为未来评估方法提供重要参考。
原文摘要 · Abstract (English)
Video-to-text summarization remains underexplored in terms of comprehensive evaluation methods. Traditional n-gram overlap-based metrics and recent large language model (LLM)-based approaches depend heavily on human-written reference summaries, limiting their practicality and sensitivity to nuanced semantic aspects. In this paper, we propose QEVA, a reference-free metric evaluating candidate summaries directly against source videos through multimodal question answering. QEVA assesses summaries along three clear dimensions: Coverage, Factuality, and Chronology. We also introduce MLVU(VS)-Eval, a new annotated benchmark derived from the MLVU dataset, comprising 800 summaries generated from 200 videos using state-of-the-art video-language multimodal models. This dataset establishes a transparent and consistent framework for evaluation. Experimental results demonstrate that QEVA shows higher correlation with human judgments compared to existing approaches, as measured by Kendall's $τ_b$, $τ_c$, and Spearman's $ρ$. We hope that our benchmark and metric will facilitate meaningful progress in video-to-text summarization research and provide valuable insights for the development of future evaluation methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。