从单张图重建真实物理交互的3D人与场景,解决深度模糊和接触不自然问题。
PhySIC: Physically Plausible 3D Human-Scene Interaction and Contact from a Single Image

- 基于图像生成带接触关系的可度量人体与场景网格
- 将场景误差降至227毫米,接触识别准确率提升至0.51
- 适合虚拟现实、机器人交互等需要真实物理约束的场景
从单张图像重建精确度量的人体与周围场景对,对虚拟现实、机器人和三维场景理解至关重要。然而现有方法在深度模糊、遮挡和物理不一致接触方面存在困难。为此,我们提出PhySIC框架,实现从单张RGB图像中联合恢复度量一致的SMPL-X人体网格、密集场景表面及顶点级接触图。该方法从粗略的单目深度和人体估计出发,通过感知遮挡的修复、可见深度与未缩放几何融合构建鲁棒度量骨架,并合成缺失支撑面如地板。通过置信加权优化,联合施加深度对齐、接触先验、穿插避免和2D重投影一致性,精修姿态、相机参数与全局尺度。显式遮挡掩码保护不可见区域免受不合理配置影响。PhySIC高效,联合优化仅需9秒,端到端不足27秒,支持多人,能自然处理复杂互动。实验表明,其将平均顶点场景误差从641毫米降至227毫米,PA-MPJPE减半至42毫米,接触F1值由0.09升至0.51。定性结果展示真实脚地交互、自然坐姿及被严重遮挡家具的合理重建。本工作将单图转化为物理合理的3D人-场景对,推动可扩展3D场景理解。代码已公开于https://yuxuan-xue.com/physic。
原文摘要 · Abstract (English)
Reconstructing metrically accurate humans and their surrounding scenes from a single image is crucial for virtual reality, robotics, and comprehensive 3D scene understanding. However, existing methods struggle with depth ambiguity, occlusions, and physically inconsistent contacts. To address these challenges, we introduce PhySIC, a framework for physically plausible Human-Scene Interaction and Contact reconstruction. PhySIC recovers metrically consistent SMPL-X human meshes, dense scene surfaces, and vertex-level contact maps within a shared coordinate frame from a single RGB image. Starting from coarse monocular depth and body estimates, PhySIC performs occlusion-aware inpainting, fuses visible depth with unscaled geometry for a robust metric scaffold, and synthesizes missing support surfaces like floors. A confidence-weighted optimization refines body pose, camera parameters, and global scale by jointly enforcing depth alignment, contact priors, interpenetration avoidance, and 2D reprojection consistency. Explicit occlusion masking safeguards invisible regions against implausible configurations. PhySIC is efficient, requiring only 9 seconds for joint human-scene optimization and under 27 seconds end-to-end. It naturally handles multiple humans, enabling reconstruction of diverse interactions. Empirically, PhySIC outperforms single-image baselines, reducing mean per-vertex scene error from 641 mm to 227 mm, halving PA-MPJPE to 42 mm, and improving contact F1 from 0.09 to 0.51. Qualitative results show realistic foot-floor interactions, natural seating, and plausible reconstructions of heavily occluded furniture. By converting a single image into a physically plausible 3D human-scene pair, PhySIC advances scalable 3D scene understanding. Our implementation is publicly available at https://yuxuan-xue.com/physic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。