arXiv:2605.15458cs.CV2026-05被引 4

用可验证奖励让视频模型学会按规则推理,提升逻辑一致性。

Video Models Can Reason with Verifiable Rewards

论文配图:Video Models Can Reason with Verifiable Rewards
图 1 · 摘自论文原文
  • 基于规则反馈优化视频生成,聚焦早期去噪阶段提升效率。
  • 在迷宫、推箱子等任务中成功率显著高于监督微调基线。
  • 适合需要严格空间/时间逻辑的视频生成场景,如仿真与交互设计。

视频扩散模型在感知真实性和时间连贯性上进展迅速,但主要仍以生成合理性为目标,缺乏可验证的推理能力。这一局限在需满足明确空间、时间或逻辑约束的任务中尤为突出。受推理导向语言模型中可验证强化学习(RLVR)的启发,我们提出VideoRLVR,一种针对视频扩散模型的实用优化方法,采用基于规则的反馈机制。VideoRLVR将视频推理建模为可验证的视觉轨迹生成,包含SDE-GRPO优化框架、密集分解奖励和早阶段聚焦策略,后者将策略优化限制在早期去噪阶段,使训练延迟降低约40%且性能不变。我们在Maze、FlowFree和Sokoban三个具有客观成功标准的程序化生成领域进行评估。结果表明,VideoRLVR在各项任务中均优于监督微调基线,尤其在低成功率场景中,密集分解奖励作用显著。其优化模型还在这些可验证推理基准及跨域测试中超越多个开源与专有视频生成模型。结果表明,可验证强化学习可推动视频模型从感知模仿迈向更可靠的规则一致视觉推理。

原文摘要 · Abstract (English)

Video diffusion models have made rapid progress in perceptual realism and temporal coherence, but they remain primarily optimized for plausible generation rather than verifiable reasoning. This limitation is especially pronounced in tasks where generated videos must satisfy explicit spatial, temporal, or logical constraints. Inspired by the role of reinforcement learning with verifiable rewards (RLVR) in reasoning-oriented language models, we introduce VideoRLVR, a practical recipe for optimizing video diffusion models with rule-based feedback. VideoRLVR formulates video reasoning as the generation of verifiable visual trajectories and consists of an SDE-GRPO optimization backbone, dense decomposed rewards, and an Early-Step Focus strategy for efficient training. The Early-Step Focus strategy restricts policy optimization to the early denoising phase, reducing training latency by about 40% while preserving performance. We evaluate VideoRLVR on Maze, FlowFree, and Sokoban, three procedurally generated domains with objective success criteria. Across these tasks, VideoRLVR consistently improves over supervised fine-tuning baselines, with dense decomposed rewards proving especially important in low-success-rate settings. Our RL-optimized model also outperforms the evaluated proprietary and open-source video generation models on these verifiable reasoning benchmarks and out-of-domain benchmarks. These results suggest that verifiable RL can move video models beyond perceptual imitation toward more reliable rule-consistent visual reasoning.

视频生成强化学习可验证推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。