arXiv:2604.08945cs.CVcs.RO2026-04被引 1

用视觉扩散模型指导稀疏触觉数据重建3D形状

TouchAnything: Diffusion-Guided 3D Reconstruction from Sparse Robot Touches

  • 用预训练视觉扩散模型提供几何先验,指导触觉重建
  • 仅需少量触点即恢复准确3D结构,优于现有方法
  • 适合未见过物体的开放世界三维重建任务

精确的物体几何估计对机器人操作和物理交互等下游任务至关重要。尽管视觉是主流形状感知模态,但在遮挡或光照困难条件下不可靠。此时触觉传感可通过物理接触提供直接几何信息。然而,仅凭稀疏局部触点重建全局3D几何本质是欠约束问题。我们提出TouchAnything框架,利用预训练的大规模2D视觉扩散模型作为语义与几何先验,指导从稀疏触觉测量中进行3D重建。不同于以往训练特定类别重建网络或直接从触觉数据学习扩散模型的方法,我们将预训练视觉扩散模型中编码的几何知识迁移到触觉域。给定稀疏接触约束与物体粗粒度类别描述,我们将重建建模为优化问题,同时满足触觉一致性并引导解向符合扩散先验的形状。本方法仅需少数触点即可重建精确几何,超越现有基线,并实现对未见物体实例的开放世界3D重建。

原文摘要 · Abstract (English)

Accurate object geometry estimation is essential for many downstream tasks, including robotic manipulation and physical interaction. Although vision is the dominant modality for shape perception, it becomes unreliable under occlusions or challenging lighting conditions. In such scenarios, tactile sensing provides direct geometric information through physical contact. However, reconstructing global 3D geometry from sparse local touches alone is fundamentally underconstrained. We present TouchAnything, a framework that leverages a pretrained large-scale 2D vision diffusion model as a semantic and geometric prior for 3D reconstruction from sparse tactile measurements. Unlike prior work that trains category-specific reconstruction networks or learns diffusion models directly from tactile data, we transfer the geometric knowledge encoded in pretrained visual diffusion models to the tactile domain. Given sparse contact constraints and a coarse class-level description of the object, we formulate reconstruction as an optimization problem that enforces tactile consistency while guiding solutions toward shapes consistent with the diffusion prior. Our method reconstructs accurate geometries from only a few touches, outperforms existing baselines, and enables open-world 3D reconstruction of previously unseen object instances. Our project page is https://grange007.github.io/touchanything .

3D重建触觉感知扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。