用声音和视觉融合提升手部姿态与接触检测精度
Visuo-Acoustic Hand Pose and Contact Estimation
- 融合骨传导发声与麦克风,通过声波传播变化感知接触
- 在遮挡和静态接触场景下,准确率显著优于纯视觉方法
- 适合机器人操作、虚拟现实等需高精度手部交互的场景
精准估计手部姿态与手物接触事件对机器人数据采集、沉浸式虚拟环境及生物力学分析至关重要,但受视觉遮挡、接触线索微弱、纯视觉感知局限及缺乏易用触觉传感等因素制约。为此,我们提出VibeMesh——一种新型可穿戴系统,结合视觉与主动声学传感,实现每顶点级的手部接触与姿态估计。该系统在人手上集成骨传导扬声器与稀疏压电麦克风,发射结构化声信号并捕捉其传播变化,以推断接触引起的扰动。为解析跨模态信号,我们设计基于图的注意力网络,同步处理音频频谱与RGB-D生成的手部网格,实现高空间分辨率接触预测。贡献包括:(i) 轻量、非侵入式多模态传感平台;(ii) 用于联合姿态与接触推断的跨模态图网络;(iii) 包含多样化操作场景下同步的RGB-D、声学与真实接触标注的数据集;(iv) 实验表明,VibeMesh在遮挡或静态接触条件下,准确率与鲁棒性均优于纯视觉基线。
原文摘要 · Abstract (English)
Accurately estimating hand pose and hand-object contact events is essential for robot data-collection, immersive virtual environments, and biomechanical analysis, yet remains challenging due to visual occlusion, subtle contact cues, limitations in vision-only sensing, and the lack of accessible and flexible tactile sensing. We therefore introduce VibeMesh, a novel wearable system that fuses vision with active acoustic sensing for dense, per-vertex hand contact and pose estimation. VibeMesh integrates a bone-conduction speaker and sparse piezoelectric microphones, distributed on a human hand, emitting structured acoustic signals and capturing their propagation to infer changes induced by contact. To interpret these cross-modal signals, we propose a graph-based attention network that processes synchronized audio spectra and RGB-D-derived hand meshes to predict contact with high spatial resolution. We contribute: (i) a lightweight, non-intrusive visuo-acoustic sensing platform; (ii) a cross-modal graph network for joint pose and contact inference; (iii) a dataset of synchronized RGB-D, acoustic, and ground-truth contact annotations across diverse manipulation scenarios; and (iv) empirical results showing that VibeMesh outperforms vision-only baselines in accuracy and robustness, particularly in occluded or static-contact settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。