arXiv:2601.03434cs.LG2026-01

首个面向多源新闻视频理解的基准数据集,测试模型跨视频对比分析能力。

VNU-Bench: A Benchmarking Dataset for Multi-Source Multimodal News Video Understanding

  • 设计新题型,评估模型跨源多模态信息对齐与整合能力
  • 构建429组新闻、1405个视频、2501个高质量问题的混合生成数据集
  • 适合研究多源信息融合、新闻推理与多模态模型评估的学者使用

新闻视频是精心剪辑的多模态叙事,融合旁白、画面和外部引述形成连贯故事。近年来,多模态大模型(MLLMs)在新闻视频理解方面取得进展,但现有基准主要关注单源、视频内推理,即每条报道独立处理。然而,真实新闻消费具有多源性:同一事件由不同媒体报导,细节互补、叙述角度各异,甚至存在冲突观点,随时间展开。因此,可靠的新闻理解需模型能比较不同来源视角,对齐跨源多模态证据,并整合多源信息。为此,我们提出VNU-Bench,首个面向新闻领域多源跨视频理解的基准。设计一系列新颖问题类型,从多角度测试模型对多源多模态新闻的理解能力。采用新型人机协同问答生成流程,解决跨源新闻理解数据集构建中的可扩展性与质量控制难题。数据集包含429组新闻、1,405个视频和2,501个高质量问题。对闭源与开源多模态模型的全面评估表明,VNU-Bench对当前MLLMs构成重大挑战。

原文摘要 · Abstract (English)

News videos are carefully edited multimodal narratives that combine narration, visuals, and external quotations into coherent storylines. In recent years, there have been significant advances in evaluating multimodal large language models (MLLMs) for news video understanding. However, existing benchmarks largely focus on single-source, intra-video reasoning, where each report is processed in isolation. In contrast, real-world news consumption is inherently multi-sourced: the same event is reported by different outlets with complementary details, distinct narrative choices, and sometimes conflicting claims that unfold over time. Robust news understanding, therefore, requires models to compare perspectives from different sources, align multimodal evidence across sources, and synthesize multi-source information. To fill this gap, we introduce VNU-Bench, the first benchmark for multi-source, cross-video understanding in the news domain. We design a set of new question types that are unique in testing models' ability of understanding multi-source multimodal news from a variety of different angles. We design a novel hybrid human-model QA generation process that addresses the issues of scalability and quality control in building a large dataset for cross-source news understanding. The dataset comprises 429 news groups, 1,405 videos, and 2,501 high-quality questions. Comprehensive evaluation of both closed- and open-source multimodal models shows that VNU-Bench poses substantial challenges for current MLLMs.

新闻理解多源分析多模态评测基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。