arXiv:2506.17685cs.CV2025-06中稿 · Pattern Recognitio…被引 4

通过动作序列提升第一视角动作识别模型在未知环境下的泛化能力。

Domain Generalization using Action Sequences for Egocentric Action Recognition

  • 利用视觉与文本上下文重建动作序列,捕捉跨域一致的用户意图。
  • 在未见环境中实现2.4%的相对性能提升,EGTEA上超越当前最优0.6%准确率。
  • 适合需要鲁棒性动作识别的机器人、可穿戴设备等实际应用场景。

从视觉输入中识别人类活动,尤其是第一人称视角,对使机器人复现人类行为至关重要。第一人称视觉通过佩戴式相机捕捉光照、视角和环境的多样化变化,导致模型在训练时未见环境中的性能显著下降。本文提出一种针对第一人称动作识别的领域泛化方法(SeqDG)。核心思想是:动作序列常反映跨视觉域一致的用户意图。为此,我们引入视觉-文本序列重建目标(SeqRec),利用文本与视觉上下文重构序列中心动作;同时通过混合不同领域的动作序列进行训练(SeqMix),增强模型鲁棒性。在EGTEA和EPIC-KITCHENS-100数据集上验证,结果表明,在未见环境中,序列表征带来2.4%的相对平均性能提升;在EGTEA上,模型达到0.6%的Top-1准确率超越当前最优水平。

原文摘要 · Abstract (English)

Recognizing human activities from visual inputs, particularly through a first-person viewpoint, is essential for enabling robots to replicate human behavior. Egocentric vision, characterized by cameras worn by observers, captures diverse changes in illumination, viewpoint, and environment. This variability leads to a notable drop in the performance of Egocentric Action Recognition models when tested in environments not seen during training. In this paper, we tackle these challenges by proposing a domain generalization approach for Egocentric Action Recognition. Our insight is that action sequences often reflect consistent user intent across visual domains. By leveraging action sequences, we aim to enhance the model's generalization ability across unseen environments. Our proposed method, named SeqDG, introduces a visual-text sequence reconstruction objective (SeqRec) that uses contextual cues from both text and visual inputs to reconstruct the central action of the sequence. Additionally, we enhance the model's robustness by training it on mixed sequences of actions from different domains (SeqMix). We validate SeqDG on the EGTEA and EPIC-KITCHENS-100 datasets. Results on EPIC-KITCHENS-100, show that SeqDG leads to +2.4% relative average improvement in cross-domain action recognition in unseen environments, and on EGTEA the model achieved +0.6% Top-1 accuracy over SOTA in intra-domain action recognition.

动作识别领域泛化第一人称视觉序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。