精准检测第一人称视频中手触物体的瞬间,提升交互与机器人学习性能。
Detecting Precise Hand Touch Moments in Egocentric Video

- 利用手部区域与上下文的跨注意力机制捕捉接触前后的细微动作变化。
- 在严格两帧容差下,平均精度比现有方法高出16.91%。
- 适用于增强现实、辅助技术及机器人学习等需要精确触碰感知的场景。
本文针对第一人称视频中精确检测手与物体接触时刻的挑战性任务展开研究。该帧级检测对增强现实、人机交互、辅助技术及机器人学习至关重要,因接触起始时刻标志着动作的开始或结束。由于接触附近手部运动细微、频繁遮挡、操作模式精细以及第一人称视角的固有运动动态,实现时间上精确的检测极具挑战。为此,我们提出手部引导的上下文增强模块(HiCE),通过手部区域及其周围上下文的时空特征,利用交叉注意力机制学习潜在接触模式。模型进一步采用抓取感知损失和软标签,强调触摸事件特有的手部姿态与运动动态,以区分近接触与实际接触帧。我们还构建了名为TouchMoment的第一人称视频数据集,包含4,021个视频和8,456个标注的接触时刻,覆盖超一百万帧。在TouchMoment上的实验表明,在严格评估标准下(仅当预测落在真实时刻前后两帧内才计为正确),我们的方法取得显著提升,平均精度优于当前最优事件定位基线16.91%。
原文摘要 · Abstract (English)
We address the challenging task of detecting the precise moment when hands make contact with objects in egocentric videos. This frame-level detection is crucial for augmented reality, human-computer interaction, assistive technologies, and robot learning applications, where contact onset signals action initiation or completion. Temporally precise detection is particularly challenging due to subtle hand motion variations near contact, frequent occlusions, fine-grained manipulation patterns, and the inherent motion dynamics of first-person perspectives. To tackle these challenges, we propose a Hand-informed Context Enhanced module (HiCE; pronounced `high-see') that leverages spatiotemporal features from hand regions and their surrounding context through cross-attention mechanisms, learning to identify potential contact patterns. Our approach is further refined with a grasp-aware loss and soft label that emphasizes hand pose patterns and movement dynamics characteristic of touch events, enabling the model to distinguish between near-contact and actual contact frames. We also introduce TouchMoment, an egocentric dataset containing 4,021 videos and 8,456 annotated contact moments spanning over one million frames. Experiments on TouchMoment show that, under a strict evaluation criterion that counts a prediction as correct only if it falls within a two-frame tolerance of the ground-truth moment, our method achieves substantial gains and outperforms state-of-the-art event-spotting baselines by 16.91% average precision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。