arXiv:2601.11087cs.CV2026-01被引 11

让视频生成模型学会遵守物理规律,自动生成更真实的物体碰撞效果。

PhysRVG: Physics-Aware Unified Reinforcement Learning for Video Generative Models

  • 用强化学习直接在高维空间强制执行物理碰撞规则,而非仅作为约束条件。
  • 提出统一框架MDcycle,实现大幅微调同时保留物理反馈能力。
  • 构建新基准PhysRVGBench,验证生成视频的物理真实性显著提升。

物理原理是实现真实视觉模拟的基础,但在基于Transformer的视频生成模型中仍被严重忽视。这一缺陷导致刚体运动渲染不准确,而刚体运动正是经典力学的核心。尽管计算机图形学和物理模拟器可轻松通过牛顿公式建模碰撞,但现代预训练-微调范式在像素级全局去噪过程中丢弃了物体刚性概念。即使数学上完全正确的约束条件,在后训练优化中也被视为次优解(即条件),从根本上限制了生成视频的物理真实性。为此,我们首次提出面向视频生成模型的物理感知强化学习范式,直接在高维空间施加物理碰撞规则,确保物理知识被严格应用而非仅作为条件。随后,我们将该范式扩展为统一框架——拟真-发现循环(MDcycle),支持大规模微调的同时,完全保留模型利用物理基反馈的能力。为验证方法有效性,我们构建了新基准PhysRVGBench,并开展广泛定性和定量实验。

原文摘要 · Abstract (English)

Physical principles are fundamental to realistic visual simulation, but remain a significant oversight in transformer-based video generation. This gap highlights a critical limitation in rendering rigid body motion, a core tenet of classical mechanics. While computer graphics and physics-based simulators can easily model such collisions using Newton formulas, modern pretrain-finetune paradigms discard the concept of object rigidity during pixel-level global denoising. Even perfectly correct mathematical constraints are treated as suboptimal solutions (i.e., conditions) during model optimization in post-training, fundamentally limiting the physical realism of generated videos. Motivated by these considerations, we introduce, for the first time, a physics-aware reinforcement learning paradigm for video generation models that enforces physical collision rules directly in high-dimensional spaces, ensuring the physics knowledge is strictly applied rather than treated as conditions. Subsequently, we extend this paradigm to a unified framework, termed Mimicry-Discovery Cycle (MDcycle), which allows substantial fine-tuning while fully preserving the model's ability to leverage physics-grounded feedback. To validate our approach, we construct new benchmark PhysRVGBench and perform extensive qualitative and quantitative experiments to thoroughly assess its effectiveness.

视频生成物理模拟强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。