用语言描述场景,让模型更准预测第一视角注意力焦点。
Robust Egocentric Visual Attention Prediction Through Language-guided Scene Context-aware Learning
- 用语言引导的上下文感知模块,生成带场景理解的视频表征。
- 在Ego4D和AEA数据集上达到最新最好效果,鲁棒性显著提升。
- 适合做第一人称视觉分析、智能辅助系统的研究者参考。
随着对第一人称视频分析需求的增长,预测相机佩戴者注意力位置(即第一人称视觉注意力预测)受到越来越多关注。然而,由于动态第一人称场景固有的复杂性和模糊性,该任务仍具挑战。受场景上下文信息对人类注意力有调节作用的启发,本文提出一种语言引导的场景上下文感知学习框架,以实现鲁棒的第一人称视觉注意力预测。首先设计了一个上下文感知器,基于语言描述的场景信息对第一人称视频进行摘要,生成上下文感知的视频表征;随后引入两种训练目标:1)促使模型聚焦于目标兴趣区域;2)抑制无关区域带来的干扰,这些区域不太可能吸引第一人称注意力。在Ego4D和Aria Everyday Activities(AEA)数据集上的大量实验表明,该方法在多样化、动态的第一人称场景中均表现出色,达到当前最优性能并具备更强鲁棒性。
原文摘要 · Abstract (English)
As the demand for analyzing egocentric videos grows, egocentric visual attention prediction, anticipating where a camera wearer will attend, has garnered increasing attention. However, it remains challenging due to the inherent complexity and ambiguity of dynamic egocentric scenes. Motivated by evidence that scene contextual information plays a crucial role in modulating human attention, in this paper, we present a language-guided scene context-aware learning framework for robust egocentric visual attention prediction. We first design a context perceiver which is guided to summarize the egocentric video based on a language-based scene description, generating context-aware video representations. We then introduce two training objectives that: 1) encourage the framework to focus on the target point-of-interest regions and 2) suppress distractions from irrelevant regions which are less likely to attract first-person attention. Extensive experiments on Ego4D and Aria Everyday Activities (AEA) datasets demonstrate the effectiveness of our approach, achieving state-of-the-art performance and enhanced robustness across diverse, dynamic egocentric scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。