提升第一视角视频定位精度,融合物体信息与拍摄者注意力机制
Object-Shot Enhanced Grounding Network for Egocentric Video
- 引入物体级信息增强视频表征,捕捉查询中强调但未被直接捕获的物体
- 利用第一视角常见镜头运动提取佩戴者注意力,改善跨模态对齐
- 在三个数据集上达到当前最优效果,适合智能穿戴应用研究者
第一视角视频定位是具身智能应用的关键任务,与外部视角视频定位存在本质差异。现有方法主要关注第一视角与外部视角视频的分布差异,却常忽视第一视角视频的关键特征及问题类型所强调的细粒度信息。为此,我们提出 OSGNet:一种面向第一视角视频的对象-视角增强定位网络。具体地,从视频中提取物体信息以丰富视频表征,尤其针对文本查询中提及但视频特征未直接包含的物体。同时,分析第一视角视频中常见的镜头运动,利用这些特征提取佩戴者的注意力信息,从而增强模型的跨模态对齐能力。在三个数据集上的实验表明,OSGNet 达到当前最优性能,验证了该方法的有效性。代码已开源:https://github.com/Yisen-Feng/OSGNet。
原文摘要 · Abstract (English)
Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentric and exocentric videos but often neglect key characteristics of egocentric videos and the fine-grained information emphasized by question-type queries. To address these limitations, we propose OSGNet, an Object-Shot enhanced Grounding Network for egocentric video. Specifically, we extract object information from videos to enrich video representation, particularly for objects highlighted in the textual query but not directly captured in the video features. Additionally, we analyze the frequent shot movements inherent to egocentric videos, leveraging these features to extract the wearer's attention information, which enhances the model's ability to perform modality alignment. Experiments conducted on three datasets demonstrate that OSGNet achieves state-of-the-art performance, validating the effectiveness of our approach. Our code can be found at https://github.com/Yisen-Feng/OSGNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。