用双手互动信息指导3D物体重建,提升遮挡下的精度与真实感。
Follow My Hold: Hand-Object Interaction Reconstruction through Geometric Guidance
- 通过手物交互几何约束,在扩散过程中优化物体形状
- 在遮挡下仍能生成高保真、物理合理的3D物体结构
- 适合需要真实手物交互的虚拟现实与机器人应用
我们提出一种基于扩散模型的新框架,仅用单目RGB图像即可重建手持物体的3D几何结构,核心是利用手物交互作为几何引导。方法在潜在空间中对物体外观进行补全,并通过推理时的闭环优化引导扩散过程,同时确保手物交互的合理性。具体而言,通过监督速度场并联合优化手和物体的变换,引入多模态几何线索:法线与深度对齐、轮廓一致性及2D关键点投影。进一步引入符号距离场监督,并强制接触与非穿透约束,保障交互的物理合理性。该方法在遮挡条件下仍能生成准确、鲁棒且一致的重建结果,且在真实场景中具有良好泛化能力。
原文摘要 · Abstract (English)
We propose a novel diffusion-based framework for reconstructing 3D geometry of hand-held objects from monocular RGB images by leveraging hand-object interaction as geometric guidance. Our method conditions a latent diffusion model on an inpainted object appearance and uses inference-time guidance to optimize the object reconstruction, while simultaneously ensuring plausible hand-object interactions. Unlike prior methods that rely on extensive post-processing or produce low-quality reconstructions, our approach directly generates high-quality object geometry during the diffusion process by introducing guidance with an optimization-in-the-loop design. Specifically, we guide the diffusion model by applying supervision to the velocity field while simultaneously optimizing the transformations of both the hand and the object being reconstructed. This optimization is driven by multi-modal geometric cues, including normal and depth alignment, silhouette consistency, and 2D keypoint reprojection. We further incorporate signed distance field supervision and enforce contact and non-intersection constraints to ensure physical plausibility of hand-object interaction. Our method yields accurate, robust and coherent reconstructions under occlusion while generalizing well to in-the-wild scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。