arXiv:2603.16870cs.CVcs.AI2026-03被引 7

发现视频模型推理本质在去噪步骤,而非逐帧链式推理。

Demystifying Video Reasoning

  • 推理发生在去噪过程中,而非按帧顺序进行。
  • 早期去噪阶段探索多解,后期逐步收敛至答案。
  • 适合研究视频生成与模型内在推理机制的学者。

近期视频生成进展揭示了一个意外现象:基于扩散的视频模型展现出非平凡的推理能力。以往研究归因于帧链(Chain-of-Frames, CoF)机制,认为推理沿视频帧顺序展开。本文挑战这一假设,揭示推理实际主要发生在扩散去噪步骤中。通过定性分析与定向探测实验,我们发现模型在早期去噪步骤中探索多个候选解,随后逐步收敛至最终答案,该过程称为链式步骤(Chain-of-Steps, CoS)。此外,我们识别出若干关键涌现推理行为:(1) 工作记忆,支持物体恒常性等需一致参照的任务;(2) 自我修正与增强,可从错误中间解中恢复;(3) 先感知后行动,早期步骤建立语义基础,后期执行结构化操作。对扩散变压器层的分析表明,中层承担关键推理任务。受此启发,我们提出一种无需训练的集成方法(Training-Free Ensemble, TFE),通过集成相同模型不同随机种子的潜在轨迹,验证了推理能力的提升。本工作首次系统解析视频推理机制,为未来利用视频模型内在推理动态构建智能提供了新范式。

原文摘要 · Abstract (English)

Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabilities. Prior work attributes this to a Chain-of-Frames (CoF) mechanism, where reasoning is assumed to unfold sequentially across video frames. In this work, we challenge this assumption and uncover a fundamentally different mechanism. We show that reasoning in video models instead primarily emerges along the diffusion denoising steps. Through qualitative analysis and targeted probing experiments, we find that models explore multiple candidate solutions in early denoising steps and progressively converge to a final answer, a process we term Chain-of-Steps (CoS). Beyond this core mechanism, we identify several emergent reasoning behaviors critical to model performance: (1) working memory that supports tasks requiring consistent reference, such as object permanence; (2) self-correction and enhancement, allowing recovery from incorrect intermediate solutions; and (3) perception before action, where early steps establish semantic grounding and later steps perform structured manipulation. Moreover, analysis of Diffusion Transformer layers shows that middle layers conduct key reasoning procedures. Motivated by these insights, we present a simple Training-Free Ensemble (TFE) as a proof-of-concept, demonstrating how reasoning can be improved by ensembling latent trajectories from identical models with different random seeds. Overall, our work provides the first systematic dissection of the mechanisms underlying video reasoning, offering a foundation to guide future research in better exploiting the inherent reasoning dynamics of video models as a new substrate for intelligence.

视频推理扩散模型去噪机制智能生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。