从单张图片重建物理稳定的3D场景,解决物体漂浮穿透问题。
REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image

- 基于重力支撑视角构建场景树,捕捉物体状态与关系
- 物理约束优化使真实场景中物理错误减少67%,模拟更稳定
- 适合虚拟现实交互与数字内容生成,可还原真实物理行为
从单张RGB图像重建物理稳定的3D场景,可将普通照片转为可用于沉浸式交互和内容创作的仿真可用数字资产。现有方法虽能生成几何合理但物理不一致的结果(如物体漂浮、穿透),导致仿真不稳定。图像条件下的场景生成方法虽提升物理合理性,却依赖强先验,造成与输入图像不符的物体布局。本文提出REST3D,通过融合物理场景理解与物理约束优化,实现物理稳定的重建。首先提出一种代理式物理场景理解技术,从重力支撑视角构建场景树表示,捕捉物体物理状态与相互关系,作为结构先验;随后利用图像到3D模型初始化场景,结合场景树引导对齐与物理约束优化,修复物理违规同时保持视觉一致性。实验表明,该方法在合成与真实数据集上显著降低物理错误(降幅达67%),提升仿真稳定性,同时保持高重建质量。进一步在基于VR的人-物交互中验证了其在沉浸式应用中的潜力。
原文摘要 · Abstract (English)
Reconstructing physically stable 3D scenes from a single RGB image enables casual images to be converted into simulation-ready digital assets for applications such as immersive interaction and content creation. However, existing single-image reconstruction methods fall short in capturing the physical structure of a scene. As a result, they often produce geometrically plausible but physically inconsistent results, including object floating and penetration, which lead to unstable behavior in physics simulations. Image-conditioned scene generation methods improve physical plausibility but often rely on strong scene priors, yielding plausible yet inaccurate object arrangements that fail to match the input image. We propose REST3D, a single-image reconstruction framework that can reconstruct physically stable 3D scenes by integrating physical scene understanding with physics-constrained refinement. We first introduce an agentic physical scene understanding technique that constructs a scene-tree representation capturing object physical states and inter-object relationships from a gravity-support perspective, providing a structural prior for reconstruction. Leveraging this structure, we initialize the scene using image-to-3D models, followed by scene-tree-guided alignment and physics-constrained optimization to resolve physical violations while preserving visual consistency with the input image. Experiments show that our method significantly reduces physical errors and improves simulation stability on both synthetic and real-world datasets while maintaining strong reconstruction quality. We further demonstrate the reconstructed scenes in VR-based human-object interaction, showing their potential for immersive applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。