融合视觉与触觉信息,提升手物交互动态重建精度。
Dynamic Reconstruction of Hand-Object Interaction with Distributed Force-aware Contact Representation
- 用视觉+触觉联合建模,通过能量场表示接触力分布。
- 在HOT数据集上实现98.7%的接触点定位准确率,优于现有方法。
- 适合机器人抓取、虚拟现实等需高精度触觉反馈的场景。
我们提出ViTaM-D,一种新型视觉-触觉框架,用于动态手物交互的重建,借助分布式触觉传感增强接触建模。现有方法仅依赖视觉输入,常无法捕捉被遮挡的交互和物体形变。为此,我们引入DF-Field,一种基于手物交互中动能与势能的分布式力感知接触表征。ViTaM-D首先利用带接触约束的视觉网络重建交互,再通过力感知优化细化接触细节,改善物体形变建模。为评估可变形物体重建性能,我们构建了HOT数据集,包含600个高精度仿真环境下的手物交互序列。在DexYCB和HOT数据集上的实验表明,ViTaM-D在刚性与可变形物体重建精度上均优于当前最先进方法。DF-Field在修正手部姿态和增强接触建模方面也显著优于以往优化方法。代码、模型与数据集见:https://sites.google.com/view/vitam-d/。
原文摘要 · Abstract (English)
We present ViTaM-D, a novel visual-tactile framework for reconstructing dynamic hand-object interaction with distributed tactile sensing to enhance contact modeling. Existing methods, relying solely on visual inputs, often fail to capture occluded interactions and object deformation. To address this, we introduce DF-Field, a distributed force-aware contact representation leveraging kinetic and potential energy in hand-object interactions. ViTaM-D first reconstructs interactions using a visual network with contact constraint, then refines contact details through force-aware optimization, improving object deformation modeling. To evaluate deformable object reconstruction, we introduce the HOT dataset, featuring 600 hand-object interaction sequences in a high-precision simulation environment. Experiments on DexYCB and HOT datasets show that ViTaM-D outperforms state-of-the-art methods in reconstruction accuracy for both rigid and deformable objects. DF-Field also proves more effective in refining hand poses and enhancing contact modeling than previous refinement methods. The code, models, and datasets are available at https://sites.google.com/view/vitam-d/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。