通过提示学习融合动作与物体语义,提升第一人称动作识别性能。
EgoPrompt: Prompt Learning for Egocentric Action Recognition
- 构建统一提示池,实现动词与名词表示的跨组件交互。
- 在三个数据集上均达到当前最佳效果,跨数据集泛化能力显著。
- 适合关注第一人称视觉理解与提示学习应用的研究者。
随着增强现实与虚拟现实应用需求增长,第一人称动作识别成为研究热点。该任务通常分为识别行为(动词)和识别作用对象(名词)两个子任务。然而,现有方法多将两者视为独立分类任务,忽视其内在语义与上下文关联,导致表征碎片化、泛化能力不足。为此,我们提出EgoPrompt框架,基于提示学习实现第一人称动作识别。在已有提示策略基础上,构建统一提示池空间,促进两类成分表示间的交互。具体而言,先将动词与名词表示分解为细粒度模式对,再通过注意力机制融合,实现跨组件信息交互。为确保提示池的丰富性,引入新颖训练目标——多样提示池准则,从提示选择频率正则与提示知识正交化两方面优化。在Ego4D、EPIC-Kitchens和EGTEA数据集上的大量实验表明,EgoPrompt在同数据集、跨数据集及基础到新类泛化任务中均表现最优。
原文摘要 · Abstract (English)
Driven by the increasing demand for applications in augmented and virtual reality, egocentric action recognition has emerged as a prominent research area. It is typically divided into two subtasks: recognizing the performed behavior (i.e., verb component) and identifying the objects being acted upon (i.e., noun component) from the first-person perspective. However, most existing approaches treat these two components as independent classification tasks, focusing on extracting component-specific knowledge while overlooking their inherent semantic and contextual relationships, leading to fragmented representations and sub-optimal generalization capability. To address these challenges, we propose a prompt learning-based framework, EgoPrompt, to conduct the egocentric action recognition task. Building on the existing prompting strategy to capture the component-specific knowledge, we construct a Unified Prompt Pool space to establish interaction between the two types of component representations. Specifically, the component representations (from verbs and nouns) are first decomposed into fine-grained patterns with the prompt pair form. Then, these pattern-level representations are fused through an attention-based mechanism to facilitate cross-component interaction. To ensure the prompt pool is informative, we further introduce a novel training objective, Diverse Pool Criteria. This objective realizes our goals from two perspectives: Prompt Selection Frequency Regularization and Prompt Knowledge Orthogonalization. Extensive experiments are conducted on the Ego4D, EPIC-Kitchens, and EGTEA datasets. The results consistently show that EgoPrompt achieves state-of-the-art performance across within-dataset, cross-dataset, and base-to-novel generalization benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。