基于早期融合的视频定位模型,斩获三类视觉记忆任务冠军
OSGNet @ Ego4D Episodic Memory Challenge 2025
- 采用早期融合策略统一处理三类定位任务
- 在自然语言、目标步骤和时间片段查询中均排名第一
- 适合需要高精度视频定位的场景应用
本文报告了我们在CVPR 2025年Ego4D情景记忆挑战赛中,针对三种第一人称视频定位任务的冠军解决方案。所有任务均要求在未剪辑的第一人称视频中精确识别时间区间。以往统一的视频定位方法多依赖晚期融合策略,往往效果不佳。为此,我们采用基于早期融合的视频定位模型,统一应对三项任务,以提升定位精度。最终,我们的方法在自然语言查询、目标步骤和时间片段查询三个赛道均取得第一名,验证了其有效性。代码已公开于https://github.com/Yisen-Feng/OSGNet。
原文摘要 · Abstract (English)
In this report, we present our champion solutions for the three egocentric video localization tracks of the Ego4D Episodic Memory Challenge at CVPR 2025. All tracks require precise localization of the interval within an untrimmed egocentric video. Previous unified video localization approaches often rely on late fusion strategies, which tend to yield suboptimal results. To address this, we adopt an early fusion-based video localization model to tackle all three tasks, aiming to enhance localization accuracy. Ultimately, our method achieved first place in the Natural Language Queries, Goal Step, and Moment Queries tracks, demonstrating its effectiveness. Our code can be found at https://github.com/Yisen-Feng/OSGNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。