单图生成物理真实、可交互的多物体视频,支持实时预览。
PhysOmni: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction

- 统一坐标系重建全场景几何,解决物体穿插与对齐模糊问题。
- 实现高保真物理交互,生成视频在物理一致性上显著优于现有方法。
- 无需训练,支持复杂力学操控,适合影视动画与虚拟现实应用。
近期生成视频模型虽具出色视觉质量,但受限于物理一致性与可控性不足。现有视频生成方法物理控制能力弱,单图转3D方法常出现物体穿插;基于物理的场景级3D生成则存在空间错位、风格化伪影及输入数据不一致问题,难以用于真实交互视频合成。我们提出PhysOmni,一种无需训练的框架,通过整体场景级3D重建,将单张图像转化为具有物理一致性和可控制性的视频。通过在统一空间坐标系中表示完整场景几何,PhysOmni解决物体穿插与对齐歧义问题。不同于以往方法,该形式支持精确的场景级多物体交互,并引入更丰富复杂的控制类型以实现高级力学操控。通过解耦模拟与渲染,PhysOmni规避了延迟高昂的先验模型,实现实时物理交互预览的同时保持照片级视觉保真度。实验表明,PhysOmni在物理保真度、空间一致性与可控性方面显著优于现有方法。
原文摘要 · Abstract (English)
Recent generative video models achieve impressive visual quality but remain constrained by limited physical consistency and controllability. Existing video generation methods provide minimal physical control, and single-image-to-3D conversion approaches often suffer from object interpenetration. Furthermore, physics-based scene-level 3D generation methods exhibit spatial misalignment, stylized artifacts, and inconsistencies with the input data, restricting their use in realistic interactive video synthesis. We propose PhysOmni, a training-free framework that converts a single image into a physically consistent and controllable video through holistic scene-level 3D reconstruction. By rep?resenting the full scene geometry in a unified spatial coordinate system, PhysOmni resolves object penetration and alignment ambiguity. Unlike prior methods, this formulation enables accurate scene?level multi-object interactions and introduces richer, complex control types for advanced mechanics?based manipulation. By decoupling simulation from rendering, PhysOmni bypasses latency-heavy priors, achieving real-time physical interaction previews paired while preserving photorealistic visual fidelity. Experimental results demonstrate that PhysOmni substantially outperforms prior methods in physical fidelity, spatial coherence, and controllability. Project Page: https://physomni.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。