通过自监督学习,让模型从无配对的主客观视频中学会跨视角理解动作。
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
- 设计自视与跨视掩码预测,同步学习视角不变表示。
- 在4个下游任务上均显著超越现有方法,提升明显。
- 适合需要跨视角视频理解的场景,如智能穿戴、自动驾驶。
从第一人称(主视角)和第三人称(客观视角)视频中学习视角不变的表征,是提升视频理解系统泛化能力的重要方向。然而,由于主客观视角在视角、运动模式和上下文上差异显著,该领域仍研究不足。本文提出一种新的掩码主-客观建模方法——自举你的视角(BYOV),旨在同时捕捉因果时序动态与跨视角对齐。我们强调了人类动作的组合性特征对鲁棒跨视角理解的重要性。具体而言,通过自视掩码和跨视掩码预测,实现视角不变且强大的视频表征学习。实验表明,BYOV在四个下游主-客观视频任务中均显著优于现有方法,各项指标均有明显提升。代码已开源:https://github.com/park-jungin/byov。
原文摘要 · Abstract (English)
View-invariant representation learning from egocentric (first-person, ego) and exocentric (third-person, exo) videos is a promising approach toward generalizing video understanding systems across multiple viewpoints. However, this area has been underexplored due to the substantial differences in perspective, motion patterns, and context between ego and exo views. In this paper, we propose a novel masked ego-exo modeling that promotes both causal temporal dynamics and cross-view alignment, called Bootstrap Your Own Views (BYOV), for fine-grained view-invariant video representation learning from unpaired ego-exo videos. We highlight the importance of capturing the compositional nature of human actions as a basis for robust cross-view understanding. Specifically, self-view masking and cross-view masking predictions are designed to learn view-invariant and powerful representations concurrently. Experimental results demonstrate that our BYOV significantly surpasses existing approaches with notable gains across all metrics in four downstream ego-exo video tasks. The code is available at https://github.com/park-jungin/byov.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。