arXiv:2606.09243cs.CVcs.AI2026-06中稿 · ICML被引 1

从第一视角视频推断手部抓握压力,无需传感器。

EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video

论文配图:EgoTactile: Learning Grasp Pressure for Everyday Objects from Egocentric Video
图 1 · 摘自论文原文
  • 用视频+物理约束建模,推断全手接触压力分布。
  • 在多样日常物体上实现高精度压力估计,跨场景泛化能力强。
  • 适合虚拟现实与机器人抓取研究者参考。

从第一视角视频中估计全手抓握压力对沉浸式虚拟现实和机器人操作至关重要,但密集触觉感知通常依赖侵入式硬件。现有视觉方法主要针对平面表面或指尖接触,难以推广到复杂三维物体交互。为此,我们提出EgoTactile,一个包含第一视角视频与全手压力标注的基准数据集,涵盖多种日常物品,并设有裸手迁移子集以支持自然场景泛化。基于该数据集,我们首先建立EgoPressureFormer作为判别基线。进一步地,为显式处理部分观测下的不确定性,提出EgoPressureDiff——一种条件扩散框架,采用大规模预训练视频扩散主干,并结合物理信息特征修正层注入语义约束,有效推断合理接触模式并解决视觉-物理歧义。大量实验表明,该方法在基准上表现优异,且具备强野外场景迁移能力。

原文摘要 · Abstract (English)

Estimating full-hand grasp pressure from egocentric video is critical for immersive VR and robotic manipulation, yet dense tactile sensing often relies on intrusive hardware. Existing vision-based methods predominantly rely on planar surfaces or fingertip contacts, failing to generalize to complex 3D object interactions. Therefore, we introduce EgoTactile, a benchmark pairing egocentric video with full-hand pressure supervision for diverse everyday objects, incorporating a bare-hand transfer subset to enable generalization to natural scenarios. Leveraging this benchmark, we first establish EgoPressureFormer as a discriminative baseline. Beyond this, to explicitly address the uncertainty in partial observations, we propose EgoPressureDiff, a conditional diffusion framework that adapts a large-scale pre-trained video diffusion backbone. By combining rich world knowledge priors with a Physically-Informed Feature Rectification layer to inject semantic constraints, our approach effectively infers plausible contact patterns and resolves visual-physical ambiguities. Extensive experiments demonstrate that our method achieves superior performance on the benchmark and robust transferability to in-the-wild scenarios. Our project page is available at https://egotactile.github.io/.

手势估计视频生成机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。