arXiv:2603.05449cs.CVcs.AI2026-03被引 17

实时生成动作驱动的视频,让物理交互变得可预演。

RealWonder: Real-Time Physical Action-Conditioned Video Generation

  • 用物理模拟作为桥梁,将动作转化为视频模型能理解的视觉信号
  • 13.2帧/秒,480x832分辨率下实现真实交互式视频生成
  • 支持刚体、柔体、流体和颗粒物的动态模拟,适合虚实融合应用

当前视频生成模型无法模拟3D动作带来的物理效应(如受力与机械操作),因其缺乏对动作如何影响3D场景的结构理解。我们提出RealWonder,首个从单图实现动作条件视频生成的实时系统。核心思路是:不直接编码连续动作,而是通过物理模拟将其转换为视频模型可处理的视觉表示(光流与RGB)。系统包含三个模块:单图3D重建、物理模拟、仅需4步扩散的轻量化视频生成器。在480x832分辨率下达到13.2 FPS,支持刚体、可变形体、流体及颗粒物的力场、机器人操作与相机控制的交互式探索。该系统为沉浸式体验、AR/VR及机器人学习开辟新可能。代码与模型权重已公开于项目主页:https://liuwei283.github.io/RealWonder/

原文摘要 · Abstract (English)

Current video generation models cannot simulate physical consequences of 3D actions like forces and robotic manipulations, as they lack structural understanding of how actions affect 3D scenes. We present RealWonder, the first real-time system for action-conditioned video generation from a single image. Our key insight is using physics simulation as an intermediate bridge: instead of directly encoding continuous actions, we translate them through physics simulation into visual representations (optical flow and RGB) that video models can process. RealWonder integrates three components: 3D reconstruction from single images, physics simulation, and a distilled video generator requiring only 4 diffusion steps. Our system achieves 13.2 FPS at 480x832 resolution, enabling interactive exploration of forces, robot actions, and camera controls on rigid objects, deformable bodies, fluids, and granular materials. We envision RealWonder opens new opportunities to apply video models in immersive experiences, AR/VR, and robot learning. Our code and model weights are publicly available in our project website: https://liuwei283.github.io/RealWonder/

视频生成物理模拟实时系统动作控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。