arXiv:2604.03701cs.CV2026-04

构建视频数字推理诊断基准,揭示视觉语言模型的失败模式。

VidNum: Diagnosing VLM Failure Modes in Video-Grounded Numerical Reasoning

  • 设计三类任务,区分直接计数、条件计数和复合推理。
  • 最佳模型准确率仅59.8%,人类达98.2%,无开源模型超45%。
  • 发现结构化目标与动作关联推理是主要瓶颈,零样本链式思考不可靠。

视频接地的数值推理要求视觉语言模型(VLMs)在帧、动作和场景变化中识别、跟踪并整合定量证据。现有基准覆盖不全:通用VideoQA将计数作为广义任务之一,而专用基准聚焦重复计数、超长视频枚举或指令数学。我们提出VidNum,一个手工标注且独立验证的基准,包含1,167道多选题。其三个任务组区分直接与独特枚举、条件与结构化枚举,以及组合性定量推理。题目级标注进一步标记证据目标、计数结构和所需推理操作。最佳评估的VLM准确率为59.8%,人类标注者为98.2%,所有开源模型均未超过45%。分层分析显示,失败并非均匀分布:结构化目标构建和动作接地的组合推理是模型反复出现的瓶颈。零样本链式思考提示并非可靠补救:虽可修复部分错误,但破坏原有正确答案,效果因模型和任务结构而异。因此,VidNum支持超越单一综合得分的诊断分析。

原文摘要 · Abstract (English)

Video-grounded numerical reasoning requires Vision-Language Models (VLMs) to identify, track, and combine quantitative evidence across frames, actions, and scene changes. Existing benchmarks provide fragmented coverage: general VideoQA includes counting among broader tasks, while dedicated benchmarks focus on repetition counting, ultra-long-video enumeration, or instructional mathematics. We introduce VidNum, a manually curated and independently verified benchmark containing 1,167 multiple-choice questions. Its three task groups distinguish Direct and Distinct Enumeration, Conditioned and Structured Enumeration, and Compositional Quantitative Reasoning. Question-level annotations further identify the evidence target, counting structure, and required reasoning operation. The best evaluated VLM reaches 59.8% accuracy, compared with 98.2% for human annotators, and no evaluated open-weight model exceeds 45%. Stratified analyses reveal that failures are not uniformly distributed: structured target construction and action-grounded compositional reasoning form recurring bottlenecks across models. Zero-shot chain-of-thought prompting is not a reliable remedy: it recovers some errors but breaks previously correct answers, with effects that vary across models and task structures. VidNum therefore supports diagnostic analysis beyond a single aggregate score.

视频理解数字推理模型诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。