arXiv:2506.04983cs.CV2025-06

首个面向长视频文字理解的评测基准,挑战大模型长期视觉文本推理能力。

TextVidBench: A Benchmark for Long Video Scene Text Understanding

  • 构建跨领域长视频数据集,平均时长2306秒,覆盖9类场景。
  • 提出三阶段评估框架,涵盖文本定位、时间对齐与动态描述生成。
  • 含5000+细粒度问答对,适合研究长视频多模态理解的学者使用。

尽管短视频文本视觉问答(ViteVQA)任务在M4-ViteVQA等基准推动下取得进展,现有数据集仍受限于视频时长和评估范围,难以充分评估多模态大模型(MLLMs)的最新能力。为此,我们提出TextVidBench,首个专为长视频文本问答(>3分钟)设计的基准。其核心贡献包括:1)跨领域长视频覆盖:涵盖新闻、体育、游戏等9类内容,平均视频长度达2306秒,更贴近真实场景;2)三阶段评估框架:‘文本寻针→时间定位→文本动态描述’;3)高质量细粒度标注:包含超5000个问答对及详细语义标签。此外,我们提出一种高效优化范式,通过引入IT-Rope机制与时间提示工程增强时间感知,采用非均匀位置编码更好处理长序列,并在视频-文本数据上进行轻量微调。在多个公开数据集及TextVidBench上的实验表明,该基准对现有模型构成显著挑战,所提方法为提升长视频场景文本理解提供了有效思路。

原文摘要 · Abstract (English)

Despite recent progress on the short-video Text-Visual Question Answering (ViteVQA) task - largely driven by benchmarks such as M4-ViteVQA - existing datasets still suffer from limited video duration and narrow evaluation scopes, making it difficult to adequately assess the growing capabilities of powerful multimodal large language models (MLLMs). To address these limitations, we introduce TextVidBench, the first benchmark specifically designed for long-video text question answering (>3 minutes). TextVidBench makes three key contributions: 1) Cross-domain long-video coverage: Spanning 9 categories (e.g., news, sports, gaming), with an average video length of 2306 seconds, enabling more realistic evaluation of long-video understanding. 2) A three-stage evaluation framework: "Text Needle-in-Haystack -> Temporal Grounding -> Text Dynamics Captioning". 3) High-quality fine-grained annotations: Containing over 5,000 question-answer pairs with detailed semantic labeling. Furthermore, we propose an efficient paradigm for improving large models through: (i) introducing the IT-Rope mechanism and temporal prompt engineering to enhance temporal perception, (ii) adopting non-uniform positional encoding to better handle long video sequences, and (iii) applying lightweight fine-tuning on video-text data. Extensive experiments on multiple public datasets as well as TextVidBench demonstrate that our new benchmark presents significant challenges to existing models, while our proposed method offers valuable insights into improving long-video scene text understanding capabilities.

视频理解长视频多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。