arXiv:2506.07886cs.CV2025-06ICCV被引 8

提出统一模型,让相机佩戴者视角视频理解更高效准确。

EgoM2P: Egocentric Multimodal Multitask Pretraining

论文配图:EgoM2P: Egocentric Multimodal Multitask Pretraining
图 1 · 摘自论文原文
  • 用时序感知的多模态标记器处理第一人称视频
  • 在多个任务上超越专用模型,速度提升十倍
  • 适合增强现实、机器人等需要理解佩戴者视角的场景

理解第一人称视觉中的多模态信号(如RGB视频、深度图、相机位姿、视线)对增强现实、机器人和人机交互至关重要,有助于系统理解佩戴者的动作、意图与环境。然而,构建大规模第一人称多模态多任务模型面临独特挑战:数据异构性强,模态覆盖差异大;缺失模态(如视线或头戴相机轨迹)难以生成伪标签,导致监督学习难扩展;动态相机运动及第一人称视频复杂的时空结构也限制了现有多模态基础模型的应用。为此,我们设计高效的时序标记器,提出EgoM2P——一种掩码建模框架,通过时序感知的多模态标记训练通用第一人称4D理解模型。该框架支持多样第一人称感知与生成任务,包括视线预测、第一人称相机跟踪、单目深度估计,并可作为条件生成模型进行第一人称视频合成。在各项任务中,EgoM2P表现媲美或优于专用模型,且速度快一个数量级。项目将完全开源,以推动第一人称视觉研究。项目页面:https://egom2p.github.io/

原文摘要 · Abstract (English)

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the camera wearer's actions, intentions, and surrounding environment. However, building large-scale egocentric multimodal and multitask models presents unique challenges. Egocentric data are inherently heterogeneous, with large variations in modality coverage across devices and settings. Generating pseudo-labels for missing modalities, such as gaze or head-mounted camera trajectories, is often infeasible, making standard supervised learning approaches difficult to scale. Furthermore, dynamic camera motion and the complex temporal and spatial structure of first-person video pose additional challenges for the direct application of existing multimodal foundation models. To address these challenges, we introduce a set of efficient temporal tokenizers and propose EgoM2P, a masked modeling framework that learns from temporally-aware multimodal tokens to train a large, general-purpose model for egocentric 4D understanding. This unified design supports multitasking across diverse egocentric perception and synthesis tasks, including gaze prediction, egocentric camera tracking, and monocular depth estimation from egocentric video, and also serves as a generative model for conditional egocentric video synthesis. Across these tasks, EgoM2P matches or outperforms specialist models while being an order of magnitude faster. We will fully open-source EgoM2P to support the community and advance egocentric vision research. Project page: https://egom2p.github.io/.

第一人称视觉多模态预训练视频生成增强现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。