构建首个分层对齐的多模态视频理解基准,解决模型泛化能力评估难题。
HAVEN: Hierarchically Aligned Multimodal Benchmark for Unified Video Understanding

- 设计分层对齐的视频-文本标注架构,覆盖帧、镜头、视频三级粒度。
- 测试主流多模态大模型在摘要与推理任务中存在表面流畅但缺乏跨模态理解的问题。
- 适合研究视频理解、多模态推理与可解释性评估的研究者使用。
尽管多模态大语言模型(MLLMs)在标准视频任务上表现优异,但其对复杂叙事的忠实总结与推理能力仍缺乏有效评估。现有摘要基准将监督信号分散于孤立粒度(如关键帧、关键镜头或割裂的文本摘要),未能捕捉跨模态对齐的固有分层结构。为此,我们提出HAVEN,一个面向统一视频理解的分层对齐多模态基准。HAVEN首创全粒度(帧、镜头、视频级别)与全模态(视频与文本)数据架构,实现模态间显式、连续的对齐。基于此统一标注范式,我们构建涵盖摘要、时间推理、多模态定位与显著性排序的综合评估套件。对前沿MLLMs的广泛评测揭示了表层文本流畅性与深层多模态理解之间的持续差距。最终,HAVEN推动多模态系统评估超越传统问答形式,提供严谨、标准化的测试平台,助力未来可解释性、分层视频理解研究。数据集、评估套件与协议已公开发布。
原文摘要 · Abstract (English)
While Multimodal Large Language Models (MLLMs) exhibit strong performance on standard video tasks, their ability to faithfully summarize and reason over complex narratives remains poorly evaluated. Existing summarization benchmarks fragment supervision across isolated granularities, such as keyframes, key shots, or disjointed text summaries, failing to capture the inherently hierarchical structure of cross-modal alignment. To address this critical gap, we introduce HAVEN, a hierarchically aligned multimodal benchmark for unified video understanding. HAVEN pioneers a fully granular (frame, shot, and video levels) and fully multimodal (video and text) dataset architecture, complete with explicit, continuous alignment between modalities. Built upon this unified annotation paradigm, we propose a comprehensive evaluation suite spanning summarization, temporal reasoning, multimodal grounding, and saliency ranking. Extensive benchmarking of state-of-the-art MLLMs exposes a persistent gap between surface-level textual fluency and grounded multimodal understanding. Ultimately, HAVEN advances the evaluation of multimodal systems beyond traditional QA formats, offering a rigorous, standardized testbed to drive future research in interpretable, hierarchical video understanding. We publicly release the dataset, benchmark suite, and evaluation protocols.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。