arXiv:2505.13091cs.CV2025-05CVPR被引 7

用触觉图像引导3D扩散模型,实现更精准的形状重建与探索

Touch2Shape: Touch-Conditioned 3D Diffusion for Shape Exploration and Reconstruction

  • 通过触觉图像条件化扩散模型,捕捉局部三维细节
  • 触觉嵌入与融合模块提升重建精度,定量指标显著优于基线
  • 结合强化学习设计触觉探索策略,适合机器人抓取与逆向设计

扩散模型在3D生成任务中取得突破。现有3D扩散模型主要依赖图像或部分观测进行目标形状重建,虽擅长全局语义理解,但难以捕捉复杂形状的局部细节,且受限于遮挡和光照条件。为克服这些局限,我们利用触觉图像捕获局部3D信息,提出Touch2Shape模型,通过触觉条件化扩散模型实现目标形状的探索与重建。在形状重建方面,设计触觉嵌入模块以条件化扩散模型生成紧凑表征,并引入触觉-形状融合模块优化重构结果。在形状探索方面,将扩散模型与强化学习结合,基于扩散模型生成的潜在向量,通过新颖奖励设计训练触觉探索策略。实验验证了重建质量的优越性,定性与定量分析均表明性能领先;触觉探索策略进一步提升了重建效果。

原文摘要 · Abstract (English)

Diffusion models have made breakthroughs in 3D generation tasks. Current 3D diffusion models focus on reconstructing target shape from images or a set of partial observations. While excelling in global context understanding, they struggle to capture the local details of complex shapes and limited to the occlusion and lighting conditions. To overcome these limitations, we utilize tactile images to capture the local 3D information and propose a Touch2Shape model, which leverages a touch-conditioned diffusion model to explore and reconstruct the target shape from touch. For shape reconstruction, we have developed a touch embedding module to condition the diffusion model in creating a compact representation and a touch shape fusion module to refine the reconstructed shape. For shape exploration, we combine the diffusion model with reinforcement learning to train a policy. This involves using the generated latent vector from the diffusion model to guide the touch exploration policy training through a novel reward design. Experiments validate the reconstruction quality thorough both qualitatively and quantitative analysis, and our touch exploration policy further boosts reconstruction performance.

3D生成触觉感知扩散模型形状重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。