arXiv:2508.12349cs.CV2025-08TPAMI被引 1

无需标注即可精确定位第一人称视频中手物接触时刻。

EgoLoc: A Generalizable Solution for Temporal Interaction Localization in Egocentric Videos

  • 用动态手势引导采样生成高质量视觉提示。
  • 零样本实现接触/分离时间点定位,准确率显著优于现有方法。
  • 适合虚拟现实、机器人操作等需要精细交互感知的场景。

分析第一人称视觉中的手物交互有助于提升虚拟现实与增强现实应用体验及人机策略迁移。现有研究多关注交互行为模式(即“如何交互”),而对关键接触与分离时刻(即“何时交互”)的精确定位仍缺乏探索,这对混合现实沉浸体验和机器人运动规划至关重要。为此,我们提出时序交互定位(Temporal Interaction Localization, TIL)问题。部分近期工作采用语义掩码作为参考,但存在对象定位不准与场景杂乱问题;尽管当前时序动作定位(TAL)方法在识别动词-名词动作段落上表现良好,却依赖类别标注训练,且在定位手物接触/分离时刻精度有限。为此,我们提出一种新型零样本方法EgoLoc,可精准定位第一人称视频中手物接触与分离的时间戳。EgoLoc引入手势动态引导采样生成高质量视觉提示,利用视觉语言模型识别接触/分离属性、定位具体时间戳,并提供闭环反馈以进一步优化。该方法无需物体掩码与动词-名词分类体系,实现可泛化的零样本部署。在公开数据集及自建基准上的综合实验表明,EgoLoc在第一人称视频中实现了可靠的TIL性能,并有效支持多种下游应用,包括第一人称视觉理解与机器人操控任务。代码与数据将开源于https://github.com/IRMVLab/EgoLoc。

原文摘要 · Abstract (English)

Analyzing hand-object interaction in egocentric vision facilitates VR/AR applications and human-robot policy transfer. Existing research has mostly focused on modeling the behavior paradigm of interactive actions (i.e., ``how to interact''). However, the more challenging and fine-grained problem of capturing the critical moments of contact and separation between the hand and the target object (i.e., ``when to interact'') is still underexplored, which is crucial for immersive interactive experiences in mixed reality and robotic motion planning. Therefore, we formulate this problem as temporal interaction localization (TIL). Some recent works extract semantic masks as TIL references, but suffer from inaccurate object grounding and cluttered scenarios. Although current temporal action localization (TAL) methods perform well in detecting verb-noun action segments, they rely on category annotations during training and exhibit limited precision in localizing hand-object contact/separation moments. To address these issues, we propose a novel zero-shot approach dubbed EgoLoc to localize hand-object contact and separation timestamps in egocentric videos. EgoLoc introduces hand-dynamics-guided sampling to generate high-quality visual prompts. It exploits the vision-language model to identify contact/separation attributes, localize specific timestamps, and provide closed-loop feedback for further refinement. EgoLoc eliminates the need for object masks and verb-noun taxonomies, leading to generalizable zero-shot implementation. Comprehensive experiments on the public dataset and our novel benchmarks demonstrate that EgoLoc achieves plausible TIL for egocentric videos. It is also validated to effectively facilitate multiple downstream applications in egocentric vision and robotic manipulation tasks. Code and relevant data will be released at https://github.com/IRMVLab/EgoLoc.

第一人称视觉交互定位零样本学习机器人操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。