让3D场景物体自然稳住,解决漂浮穿插问题
$ϕ$-Scene: Physically Grounded Image-to-3D Scene Reconstruction

- 按拓扑顺序逐个摆放物体,模拟真实物理支撑关系
- 在3D-Front上优于多数跨域方法,穿插减少90%以上
- 适合需要真实物理交互的场景重建应用
现有图像到3D场景重建方法虽能恢复高保真物体和合理布局,但常出现漂浮与穿插现象,影响物理合理性及交互环境中的使用。本文提出ϕ-Scene,一种开放词汇、可组合的物理驱动图像到3D场景重建方法,将场景视为全局稳定的物理系统而非仅是带预测位姿的物体集合。ϕ-Scene将重建建模为拓扑引导的物理组装:给定初始重建物体场景,推断物体间的支撑关系,并按拓扑顺序逐个放置。每个物体先通过SDF优化解决与已固定支撑物的穿透问题,再通过刚体物理仿真,在真实物理约束下达到稳定平衡。最终场景保持与参考图像对齐,所有物体均处于物理有效且稳定的接触状态。在3D-Front基准上,ϕ-Scene在跨域方法中表现最强,且在标准重建指标上媲美域内基线。人类与多模态大模型评估显示其在视觉质量、参考对齐和物理合理性方面更优。专用物理指标表明,其显著减少穿透伪影,仿真后漂移极低。据我们所知,ϕ-Scene是首个在保持参考对齐的同时显式达成动态刚体平衡的图像到3D场景重建方法。
原文摘要 · Abstract (English)
Recent image-to-3D scene methods recover high-fidelity 3D objects with plausible arrangements, but often leave floatings and interpenetrations that limit physical validity and downstream use in interactive environments. We present $ϕ$-Scene, a physically grounded approach for open-vocabulary and compositional image-to-3D scene reconstruction that treats a scene not merely as a set of objects with predicted poses, but as a globally stable physical system. $ϕ$-Scene formulates reconstruction as topology-driven physical assembly: given an initial scene of reconstructed objects, it infers how objects support one another and settles them one by one in topological order. For each object, SDF-based optimization first resolves penetrations against the already-settled support context, and rigid-body simulation then settles the object into a stable equilibrium under real-world physical constraints. The resulting scene stays aligned to the reference image, with every object resting at a physically valid, stable contact configuration. On the 3D-Front benchmark, $ϕ$-Scene achieves the strongest overall performance among out-of-domain methods and remains highly competitive with in-domain baselines on standard reconstruction metrics. Human and MLLM studies prefer $ϕ$-Scene in visual quality, reference alignment, and physical plausibility. Dedicated physical metrics show that it substantially reduces penetration artifacts and yields much lower post-simulation drift. To our knowledge, $ϕ$-Scene is among the first image-to-3D scene reconstruction methods that explicitly reaches dynamic rigid-body equilibrium while preserving reference alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。