评测模型在视频中跨模态数学推理能力,挑战真实场景下的复杂理解。
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
- 构建多模态视频数学题库,融合视觉、音频与文本信息。
- 涵盖10个数学领域,最长视频超1小时,需多步推理作答。
- 适合研究视频理解、多模态推理与AI教育的学者使用。
现实世界中的视频数学推理与静态图像或文本不同,需解析细粒度视觉信息、准确读取手写或数字文字,并整合分散的语音线索,常呈非线性分布。在此类多模态背景下,成功不仅依赖感知,更在于从丰富而嘈杂的内容流中选择性识别并整合关键上下文。为此,我们提出 VideoMathQA,一个用于评估模型在视频中执行时序延展的跨模态推理能力的基准。该基准覆盖10个不同的数学领域,包含时长10秒至超过1小时的视频。要求模型解读结构化视觉内容,理解教学叙述,并联合对齐视觉、音频和文本模态中的概念。我们动用研究生级专家进行标注,累计投入超920人时。为反映真实场景,问题设计围绕三大核心推理挑战:直接解题(答案基于呈现的问题)、概念迁移(将已学方法应用于新问题)以及深度教学理解(对长时间解释和部分解答进行多步推理)。每个问题附带多步推理标注,支持对模型能力的细粒度诊断。通过该基准,我们揭示现有方法的局限性,并建立系统性评估框架,以检验模型能否在时序延展且模态丰富的数学问题设置中真正推理而非仅感知。基准与评估代码已公开:https://mbzuai-oryx.github.io/VideoMathQA
原文摘要 · Abstract (English)
Mathematical reasoning in real-world video settings presents a fundamentally different challenge than in static images or text. It requires interpreting fine-grained visual information, accurately reading handwritten or digital text, and integrating spoken cues, often dispersed non-linearly over time. In such multimodal contexts, success hinges not just on perception, but on selectively identifying and integrating the right contextual details from a rich and noisy stream of content. To this end, we introduce VideoMathQA, a benchmark designed to evaluate whether models can perform such temporally extended cross-modal reasoning on videos. The benchmark spans 10 diverse mathematical domains, covering videos ranging from 10 seconds to over 1 hour. It requires models to interpret structured visual content, understand instructional narratives, and jointly ground concepts across visual, audio, and textual modalities. We employ graduate-level experts to ensure high quality, totaling over $920$ man-hours of annotation. To reflect real-world scenarios, questions are designed around three core reasoning challenges: direct problem solving, where answers are grounded in the presented question; conceptual transfer, which requires applying learned methods to new problems; and deep instructional comprehension, involving multi-step reasoning over extended explanations and partially worked-out solutions. Each question includes multi-step reasoning annotations, enabling fine-grained diagnosis of model capabilities. Through this benchmark, we highlight the limitations of existing approaches and establish a systematic evaluation framework for models that must reason, rather than merely perceive, across temporally extended and modality-rich mathematical problem settings. Our benchmark and evaluation code are available at: https://mbzuai-oryx.github.io/VideoMathQA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。