arXiv:2512.01424cs.CV2025-12被引 1

构建视频推理纠错基准,评估大模型发现并修正错误的能力。

ViRectify: A Challenging Benchmark for Video Reasoning Correction with Multimodal Large Language Models

  • 设计多阶段纠错流程,要求模型定位错误并基于视频证据生成解释。
  • 16个先进模型在3万+样本上测试,最高仅31.94%正确率,挑战巨大。
  • 适合研究视频理解、模型反思与纠错机制的学者使用。

多模态大语言模型(MLLMs)在复杂视频推理中常出现错误,纠正这些错误对揭示其弱点和提升性能至关重要。然而,现有基准缺乏对MLLMs识别与修正视频推理错误能力的系统评估。为此,我们提出ViRectify,一个全面的评测基准,用于评估模型细粒度纠错能力。通过AI辅助标注与人工验证,构建了涵盖动态感知、科学推理和具身决策领域的超3万实例数据集。在ViRectify中,模型需分步识别错误,并生成基于关键视频证据的推理过程。我们进一步提出轨迹证据驱动的纠错框架,包含分步错误轨迹建模与视觉证据支撑的奖励机制,引导模型聚焦错误传播路径与关键时间戳。对16个先进MLLM的评估表明,该基准极具挑战性,GPT-5仅达31.94%纠错准确率。我们的框架使Qwen2.5-VL-7B持续优于72B版本,验证了方法有效性。进一步分析揭示了模型间纠错行为的系统性不对称,该数据集亦可支持反思学习研究。我们认为ViRectify为全面评估高级视频推理模型提供了新方向。

原文摘要 · Abstract (English)

As multimodal large language models (MLLMs) frequently exhibit errors in complex video reasoning scenarios, correcting these errors is critical for uncovering their weaknesses and improving performance. However, existing benchmarks lack systematic evaluation of MLLMs' ability to identify and correct these video reasoning errors. To bridge this gap, we propose ViRectify, a comprehensive benchmark to evaluate their fine-grained correction capability. Through an AI-assisted annotation pipeline with human verification, we construct a dataset of over 30K instances spanning dynamic perception, scientific reasoning, and embodied decision-making domains. In ViRectify, we challenge MLLMs to perform step-wise error identification and generate rationales with key video evidence grounding. In addition, we further propose the trajectory evidence-driven correction framework, comprising step-wise error trajectory and reward modeling on visual evidence-grounded correction. It encourages the model to explicitly concentrate on error propagation and key timestamps for correction. Extensive evaluation across 16 advanced MLLMs demonstrates that our ViRectify serves as a challenging testbed, where GPT-5 achieves only 31.94% correction accuracy. Our framework enables a Qwen2.5-VL-7B to consistently outperform the variants of 72B on ViRectify, showing the effectiveness of our approach. Further analysis uncovers systematic asymmetries in error correction across models, and our dataset is also a valuable data resource to perform reflection learning. We believe ViRectify provides a new direction for comprehensively evaluating the advanced MLLMs in video reasoning.

视频推理纠错大模型基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。