arXiv:2505.22810cs.CV2025-05被引 9

构建首个面向视频文本理解的综合性评估基准,填补现有评测空白。

VidText: Towards Comprehensive Evaluation for Video Text Understanding

  • 设计分层评估框架,覆盖视频、片段和实例三级任务
  • 在18个主流多模态模型上测试,普遍表现不佳,提升空间大
  • 支持多语言与真实场景,适合研究动态环境下图文交互的学者

嵌入视频中的视觉文本包含丰富语义信息,对整体视频理解与局部人类动作细粒度推理至关重要。然而,现有视频理解基准大多忽略文本信息,而OCR专用基准仅限静态图像,难以捕捉文本与动态视觉上下文的交互。为此,我们提出VidText,一个面向视频文本理解的综合性评估基准。该基准具备三大特性:1)覆盖多种真实场景,支持多语言内容;2)引入分层评估框架,包含视频级、片段级与实例级任务,可评估全局摘要与局部检索能力;3)设计成对的感知推理任务,涵盖从视觉文本识别到跨模态推理。在18个先进多模态大模型上的实验表明,当前模型在多数任务中表现不佳,仍有巨大提升空间。进一步分析揭示了模型内在因素(如输入分辨率、OCR能力)与外部因素(如辅助信息、思维链策略)的影响。我们希望VidText能填补视频理解评测的空白,成为未来动态环境中多模态推理研究的基础。

原文摘要 · Abstract (English)

Visual texts embedded in videos carry rich semantic information, which is crucial for both holistic video understanding and fine-grained reasoning about local human actions. However, existing video understanding benchmarks largely overlook textual information, while OCR-specific benchmarks are constrained to static images, limiting their ability to capture the interaction between text and dynamic visual contexts. To address this gap, we propose VidText, a new benchmark designed for comprehensive and in-depth evaluation of video text understanding. VidText offers the following key features: 1) It covers a wide range of real-world scenarios and supports multilingual content, encompassing diverse settings where video text naturally appears. 2) It introduces a hierarchical evaluation framework with video-level, clip-level, and instance-level tasks, enabling assessment of both global summarization and local retrieval capabilities. 3) The benchmark also introduces a set of paired perception reasoning tasks, ranging from visual text perception to cross-modal reasoning between textual and visual information. Extensive experiments on 18 state-of-the-art Large Multimodal Models (LMMs) reveal that current models struggle across most tasks, with significant room for improvement. Further analysis highlights the impact of both model-intrinsic factors, such as input resolution and OCR capability, and external factors, including the use of auxiliary information and Chain-of-Thought reasoning strategies. We hope VidText will fill the current gap in video understanding benchmarks and serve as a foundation for future research on multimodal reasoning with video text in dynamic environments.

视频文本多模态评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。