arXiv:2608.29711cs.CL2026-08

评测大模型理解动态图表与叙事视频的能力,发现开源模型表现不随参数增大而提升。

DVBench: Benchmarking MLLMs for Understanding Dynamic Charts and Narratives in Data Videos

论文配图:DVBench: Benchmarking MLLMs for Understanding Dynamic Charts and Narratives in Data Videos
图 1 · 摘自论文原文
  • 构建五维评估框架,用真实数据视频和人工验证问题对齐多模态理解能力
  • 9个模型测试显示Gemini-3.1-Pro最优,开源模型中Kimi-k2.5最强
  • 揭示开源模型性能不随参数增长,叙事能力强未必视觉理解强

尽管多模态大模型在图表理解和视频理解方面取得显著进展,但现有评估大多孤立地考察这些能力,缺乏对时序演化的结构化视觉信息的理解。为此,我们提出DVBench,一个针对数据视频的评测基准,这是一种融合动态图表与结构化叙事的故事讲述媒介。我们将数据视频理解分解为五个维度。DVBench包含300个真实世界的数据视频和1000个经人工验证的问答对,通过严格的半自动化流程筛选。对九个MLLMs的广泛评估显示,Gemini-3.1-Pro整体表现最佳,而Kimi-k2.5是性能最强的开源模型。我们进一步识别出两个显著现象:开源模型性能并不严格随参数规模增长;叙事能力强者不一定具备强视觉理解能力。细粒度分析与消融实验揭示了各维度的弱点,以及帧配置和字幕输入的影响,为未来MLLM发展提供指导。DVBench已公开发布于https://bomiaowang.github.io/DVBench/。

原文摘要 · Abstract (English)

While MLLMs have made significant strides in chart comprehension and video understanding, current evaluations largely isolate these capabilities, leaving a critical gap in understanding temporally evolving structured visual information. To address this gap, we introduce DVBench, a benchmark for evaluating MLLMs on data videos, a storytelling medium that integrates dynamic charts with structured narratives. We decompose data video understanding into five dimensions. DVBench comprises 300 real-world data videos and 1,000 human-verified QA pairs curated through a rigorous semi-automated pipeline. Extensive evaluations of nine MLLMs show that Gemini-3.1-Pro achieves the best overall performance, while Kimi-k2.5 is the strongest open-source model. We further identify two notable phenomena: open-source model performance does not scale strictly with parameter size, and narrative proficiency does not guarantee visual capability. Fine-grained analyses and ablation studies further reveal dimension-specific weaknesses and the effects of frame configurations and subtitle inputs, informing future MLLM development. DVBench is publicly available at https://bomiaowang.github.io/DVBench/.

多模态评测基准动态图表视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。