分离动作区域的通用特征,提升跨场景自拍动作识别能力
Egocentric zone-aware action recognition across environments
- 将动作区域的通用语义与具体场景外观解耦
- 在EPIC-Kitchens-100和Argo1M上实现更强跨域迁移性能
- 适合研究自拍视觉与跨场景动作识别的学者
人类活动与其所处位置密切相关,如在水槽边洗手。日常环境中存在特定的活动中心区域(activity-centric zones),这些区域通常支持一类同质动作。利用这些区域知识可作为先验信息,帮助视觉模型识别动作。然而,这些区域的外观具有场景特异性,限制了先验信息在陌生环境中的泛化能力。该问题在自拍视觉中尤为突出,因环境占据图像大部分空间,导致动作与上下文难以分离。本文探讨将活动中心区域的领域特定外观与其通用、领域无关表示相解耦的重要性,并验证这种通用表示能显著提升自拍动作识别(EAR)模型的跨域迁移能力。实验在EPIC-Kitchens-100和Argo1M数据集上进行。
原文摘要 · Abstract (English)
Human activities exhibit a strong correlation between actions and the places where these are performed, such as washing something at a sink. More specifically, in daily living environments we may identify particular locations, hereinafter named activity-centric zones, which may afford a set of homogeneous actions. Their knowledge can serve as a prior to favor vision models to recognize human activities. However, the appearance of these zones is scene-specific, limiting the transferability of this prior information to unfamiliar areas and domains. This problem is particularly relevant in egocentric vision, where the environment takes up most of the image, making it even more difficult to separate the action from the context. In this paper, we discuss the importance of decoupling the domain-specific appearance of activity-centric zones from their universal, domain-agnostic representations, and show how the latter can improve the cross-domain transferability of Egocentric Action Recognition (EAR) models. We validate our solution on the EPIC-Kitchens-100 and Argo1M datasets
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。