arXiv:2603.20850cs.CVcs.RO2026-03被引 5

用传感器手套生成逼真手物交互视频,保留物理互动细节。

Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves

  • 从多模态手套数据重建逼真裸手动作,保持物理动态一致。
  • 构建首个含触觉与惯性信号的多模态手物交互数据集HandSense。
  • 适合研究手部交互、虚拟现实及机器人感知的开发者使用。

理解手物交互(HOI)是计算机视觉、机器人和AR/VR的基础。然而,传统手部视频常缺少接触力、运动信号等关键物理信息,且易受遮挡影响。为此,我们提出Glove2Hand框架,将多模态传感手套采集的HOI视频转化为逼真裸手视频,同时忠实保留底层物理交互动态。引入新型3D高斯手部模型,确保时间渲染一致性;通过基于扩散模型的手部修复器,实现复杂手物交互与非刚性形变的自然融合。基于Glove2Hand,我们构建了首个包含手套到手部视频并同步触觉与IMU信号的多模态HOI数据集HandSense。实验表明,HandSense显著提升下游裸手应用性能,包括基于视频的接触力估计与严重遮挡下的手部追踪。

原文摘要 · Abstract (English)

Understanding hand-object interaction (HOI) is fundamental to computer vision, robotics, and AR/VR. However, conventional hand videos often lack essential physical information such as contact forces and motion signals, and are prone to frequent occlusions. To address the challenges, we present Glove2Hand, a framework that translates multi-modal sensing glove HOI videos into photorealistic bare hands, while faithfully preserving the underlying physical interaction dynamics. We introduce a novel 3D Gaussian hand model that ensures temporal rendering consistency. The rendered hand is seamlessly integrated into the scene using a diffusion-based hand restorer, which effectively handles complex hand-object interactions and non-rigid deformations. Leveraging Glove2Hand, we create HandSense, the first multi-modal HOI dataset featuring glove-to-hand videos with synchronized tactile and IMU signals. We demonstrate that HandSense significantly enhances downstream bare-hand applications, including video-based contact estimation and hand tracking under severe occlusion.

手物交互多模态数据集生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。