arXiv:2604.09100cs.CVcs.RO2026-04中稿 · ECCV

用手部感知和触觉实现被手遮挡物体的物理可信三维重建

Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch

论文配图:Physically Grounded 3D Generative Reconstruction under Hand Occlusion using Proprioception and Multi-Contact Touch
图 1 · 摘自论文原文
  • 结合本体感觉与多点触觉,用物理约束减少遮挡区域的不确定性
  • 在模拟与真实机器人上均实现正确尺度、无穿透的重建结果
  • 适合需要高精度三维感知的机器人操作任务

我们提出一种多模态、物理可信的方法,用于在严重手部遮挡条件下进行度量尺度的非完整物体重建与位姿估计。不同于仅依赖视觉的已有方法,本工作利用物理交互信号:本体感觉提供手部姿态几何,多点触觉约束物体表面位置,降低遮挡区域的歧义性。我们将物体结构表示为姿态感知、相机对齐的有符号距离场(SDF),并使用结构-变分自编码器(Structure-VAE)学习紧凑隐空间。在此隐空间中,训练一个条件流匹配扩散模型,先在纯视觉图像上预训练,再在遮挡操作场景中微调,同时以可见RGB证据、遮挡物/可见性掩码、手部隐表示及触觉信息为条件。关键的是,在微调与推理过程中引入基于物理的目标函数和可微解码器引导,以减少手-物穿插,并使重建表面与接触观测对齐。由于该方法生成度量一致、物理合理的结构估计,可无缝集成至现有两阶段重建流程中,下游模块进一步优化几何并预测外观。仿真实验表明,加入本体感觉与触觉显著提升遮挡下的补全效果,生成结果在真实世界尺度下具有物理合理性;进一步通过将模型部署于实际人形机器人验证迁移能力,其末端执行器与训练时不同。

原文摘要 · Abstract (English)

We propose a multimodal, physically grounded approach for metric-scale amodal object reconstruction and pose estimation under severe hand occlusion. Unlike prior occlusion-aware 3D generation methods that rely only on vision, we leverage physical interaction signals: proprioception provides the posed hand geometry, and multi-contact touch constrains where the object surface must lie, reducing ambiguity in occluded regions. We represent object structure as a pose-aware, camera-aligned signed distance field (SDF) and learn a compact latent space with a Structure-VAE. In this latent space, we train a conditional flow-matching diffusion model, pretraining on vision-only images and finetuning on occluded manipulation scenes while conditioning on visible RGB evidence, occluder/visibility masks, the hand latent representation, and tactile information. Crucially, we incorporate physics-based objectives and differentiable decoder-guidance during finetuning and inference to reduce hand--object interpenetration and to align the reconstructed surface with contact observations. Because our method produces a metric, physically consistent structure estimate, it integrates naturally into existing two-stage reconstruction pipelines, where a downstream module refines geometry and predicts appearance. Experiments in simulation show that adding proprioception and touch substantially improves completion under occlusion and yields physically plausible reconstructions at correct real-world scale compared to vision-only baselines; we further validate transfer by deploying the model on a real humanoid robot with an end-effector different from those used during training.

三维重建触觉感知机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。