arXiv:2602.07689cs.CVcs.AI2026-02被引 2

让视频推理过程可追踪,提升准确性和可解释性

Process-of-Thought Reasoning for Videos

  • 将视频推理拆解为可验证的步骤,逐步更新假设
  • 在多个任务上显著提升事实正确率和时间定位精度
  • 适合需要透明推理的场景,如医疗或司法视频分析

视频理解不仅需识别视觉内容,还需对长时序、噪声多的观测进行时空定位的多步推理。本文提出视频领域的思维链(PoT)推理框架,通过将视频推理结构化为一系列轻量、可验证的步骤,使推理过程显式化。PoT 交织执行(i)时间证据选择,(ii)分步状态更新,(iii)受限答案生成,使模型能逐步修正假设并保持与视频证据的可追溯性。该框架不依赖特定模型,可接入现有视觉-语言主干网络,支持闭卷推理与外部工具增强推理。我们进一步提出统一的 PoT 轨迹表示,将中间决策对齐至时间片段,增强对干扰项的鲁棒性,减少幻觉解释。在标准视频推理任务上的大量实验表明,PoT 持续提升事实正确率与时间定位能力,同时提供可诊断的可解释推理轨迹。

原文摘要 · Abstract (English)

Video understanding requires not only recognizing visual content but also performing temporally grounded, multi-step reasoning over long and noisy observations. We propose Process-of-Thought (PoT) Reasoning for Videos, a framework that makes the reasoning process explicit by structuring video inference into a sequence of lightweight, verifiable steps. PoT interleaves (i) temporal evidence selection, (ii) step-wise state updates, and (iii) constrained answer synthesis, enabling the model to progressively refine hypotheses while maintaining traceability to video evidence. The framework is designed to be model-agnostic and can be plugged into existing vision-language backbones, supporting both closed-book reasoning and evidence-augmented reasoning with external tools. We further introduce a unified representation for PoT traces that aligns intermediate decisions with temporal segments, which improves robustness to distractors and reduces hallucinated explanations. Extensive experiments on standard video reasoning tasks demonstrate that PoT consistently improves factual correctness and temporal grounding, while providing interpretable reasoning traces for diagnosis and downstream use.

视频理解思维链可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。