解决第一人称视频中物体定位难题,提升复杂视角下的准确率。
HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization
- 模仿人类认知,用上下文注意力与局部特征优化定位
- 在VQ2D数据集上超越基线模型,定位精度显著提升
- 适合研究第一人称视觉理解与鲁棒匹配的学者
本文针对第一人称视觉查询定位(VQL)任务,旨在定位长时序第一人称视频中的目标物体。由于第一人称视频中频繁且剧烈的视角变化导致物体外观大幅变异和部分遮挡,现有方法难以实现精准定位。为此,我们提出一种受人类认知过程启发的新方法——层次化、第一人称且鲁棒的视觉查询定位(HERO-VQL),包含两项核心创新:(i) 上下文注意力引导(TAG),利用类别标记提供高层语义上下文,结合主成分得分图实现细粒度定位;(ii) 第一人称增强一致性训练(EgoACT),通过替换查询为真实标注中的随机对应物体来增强查询多样性,并通过重排视频帧模拟极端视角变化。同时引入一致性训练损失(CT loss),确保在不同增强场景下定位结果稳定。在VQ2D数据集上的大量实验表明,HERO-VQL能有效应对第一人称视频挑战,显著优于基线方法。
原文摘要 · Abstract (English)
In this work, we tackle the egocentric visual query localization (VQL), where a model should localize the query object in a long-form egocentric video. Frequent and abrupt viewpoint changes in egocentric videos cause significant object appearance variations and partial occlusions, making it difficult for existing methods to achieve accurate localization. To tackle these challenges, we introduce Hierarchical, Egocentric and RObust Visual Query Localization (HERO-VQL), a novel method inspired by human cognitive process in object recognition. We propose i) Top-down Attention Guidance (TAG) and ii) Egocentric Augmentation based Consistency Training (EgoACT). Top-down Attention Guidance refines the attention mechanism by leveraging the class token for high-level context and principal component score maps for fine-grained localization. To enhance learning in diverse and challenging matching scenarios, EgoAug enhances query diversity by replacing the query with a randomly selected corresponding object from groundtruth annotations and simulates extreme viewpoint changes by reordering video frames. Additionally, CT loss enforces stable object localization across different augmentation scenarios. Extensive experiments on VQ2D dataset validate that HERO-VQL effectively handles egocentric challenges, significantly outperforming baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。