arXiv:2506.05782cs.CV2025-06

利用视线信息提升自然语言视频检索准确率

GazeNLQ @ Ego4D Natural Language Queries Challenge 2025

  • 通过对比学习直接从视频中预训练视线估计模型
  • 在Ego4D数据集上达到27.82的[email protected]和18.68的[email protected]
  • 适合关注视觉注意力与视频理解结合的研究者

本报告介绍我们在CVPR 2025年Ego4D自然语言查询(NLQ)挑战赛中的解决方案。第一人称视频从佩戴者视角捕捉场景,其中视线作为关键非语言交流线索,反映视觉注意力并揭示人类意图与认知。受此启发,我们提出GazeNLQ方法,利用视线信息检索匹配自然语言查询的视频片段。具体而言,我们设计了一种基于对比学习的视线估计预训练策略,直接从视频中学习。估计出的视线用于增强所提模型中的视频表征,从而提升定位精度。实验结果表明,GazeNLQ在[email protected][email protected]指标上分别达到27.82和18.68。代码已公开于https://github.com/stevenlin510/GazeNLQ。

原文摘要 · Abstract (English)

This report presents our solution to the Ego4D Natural Language Queries (NLQ) Challenge at CVPR 2025. Egocentric video captures the scene from the wearer's perspective, where gaze serves as a key non-verbal communication cue that reflects visual attention and offer insights into human intention and cognition. Motivated by this, we propose a novel approach, GazeNLQ, which leverages gaze to retrieve video segments that match given natural language queries. Specifically, we introduce a contrastive learning-based pretraining strategy for gaze estimation directly from video. The estimated gaze is used to augment video representations within proposed model, thereby enhancing localization accuracy. Experimental results show that GazeNLQ achieves [email protected] and [email protected] scores of 27.82 and 18.68, respectively. Our code is available at https://github.com/stevenlin510/GazeNLQ.

视频检索视线估计第一人称视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。