arXiv:2502.07707cs.CV2025-02ICCV被引 8

通过视频内目标知识渐进式优化,提升第一人称视频中目标定位的鲁棒性。

PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization

  • 利用视频内目标外观与空间知识逐步精炼查询和视频特征
  • 在Ego4D数据集上达到当前最优性能,显著超越已有方法
  • 适合复杂场景下第一人称视觉定位任务的研究与应用

第一人称视觉查询定位(EgoVQL)旨在给定视觉查询的情况下,从第一人称视频中定位感兴趣的目标在时空中的位置。尽管已有方法不断进步,但面对目标外观剧烈变化和背景杂乱的情况时仍表现不佳,主要因缺乏足够的目标线索。为此,本文提出一种渐进式知识引导精炼框架PRVQL。其核心思想是持续从视频中提取与目标相关的知识,并用作指导来逐步优化查询和视频特征。框架包含多个处理阶段,前一阶段提取的外观与空间知识作为下一阶段的指导信号,用于进一步精炼特征并生成更准确的知识。该渐进过程使目标知识逐步增强,最终提升定位精度。实验表明,在具有挑战性的Ego4D数据集上,PRVQL取得当前最佳结果,显著优于现有方法。代码、模型及结果将开源。

原文摘要 · Abstract (English)

Egocentric visual query localization (EgoVQL) focuses on localizing the target of interest in space and time from first-person videos, given a visual query. Despite recent progressive, existing methods often struggle to handle severe object appearance changes and cluttering background in the video due to lacking sufficient target cues, leading to degradation. Addressing this, we introduce PRVQL, a novel Progressive knowledge-guided Refinement framework for EgoVQL. The core is to continuously exploit target-relevant knowledge directly from videos and utilize it as guidance to refine both query and video features for improving target localization. Our PRVQL contains multiple processing stages. The target knowledge from one stage, comprising appearance and spatial knowledge extracted via two specially designed knowledge learning modules, are utilized as guidance to refine the query and videos features for the next stage, which are used to generate more accurate knowledge for further feature refinement. With such a progressive process, target knowledge in PRVQL can be gradually improved, which, in turn, leads to better refined query and video features for localization in the final stage. Compared to previous methods, our PRVQL, besides the given object cues, enjoys additional crucial target information from a video as guidance to refine features, and hence enhances EgoVQL in complicated scenes. In our experiments on challenging Ego4D, PRVQL achieves state-of-the-art result and largely surpasses other methods, showing its efficacy. Our code, model and results will be released at https://github.com/fb-reps/PRVQL.

目标定位第一人称视频知识引导Ego4D

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。