提出统一评估多模态视频摘要信息损失的ViSIL指标
ViSIL: Unified Evaluation of Information Loss in Multimodal Video Captioning
- 基于信息论构建跨模态信息损失度量框架
- 与人类和VLM在VQA任务上的表现显著相关
- 可优化摘要生成速度与准确率的权衡
多模态视频字幕将密集视频内容浓缩为关键帧与自然语言的结构化摘要,为生成式AI提供语义依据,并作为高效检索的轻量代理。然而,传统指标如BLEU或ROUGE无法衡量文本与关键帧等异构模态间的信息覆盖程度。为此,本文提出视频摘要信息损失(ViSIL)评分,一种基于视觉-语言模型(VLM)推理的信息理论框架,量化摘要未能捕捉的视频信息量。通过测量信息损失,ViSIL实现了不同结构摘要间的统一比较。实验表明,ViSIL得分与人类及VLM在视频问答(VQA)任务中的表现存在显著统计相关性。此外,该指标支持摘要选择以优化信息损失与处理速度的权衡,在不增加计算负担的前提下,使摘要性能比纯文本提升7%的VQA准确率。
原文摘要 · Abstract (English)
Multimodal video captioning condenses dense footage into a structured format of keyframes and natural language. By creating a cohesive multimodal summary, this approach anchors generative AI in rich semantic evidence and serves as a lightweight proxy for high-efficiency retrieval. However, traditional metrics like BLEU or ROUGE fail to quantify information coverage across disparate modalities, such as comparing a paragraph of text to a sequence of keyframes. To address this, we propose the Video Summary Information Loss (ViSIL) score, an information-theoretic framework that quantifies the video information not captured by a summary via vision-language model (VLM) inference. By measuring the information loss, ViSIL is a unified metric that enables direct comparison across multimodal summary formats despite their structural discrepancies. Our results demonstrate that ViSIL scores show a statistically significant correlation with both human and VLM performance on Video Question Answering (VQA) tasks. ViSIL also enables summary selection to optimize the trade-off between information loss and processing speed, establishing a Pareto-optimal frontier that outperforms text summaries by $7\%$ in VQA accuracy without increasing processing load.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。