用第三人称知识提升第一人称视频理解,解决数据少、模型差的问题。
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
- 利用第三人称视觉知识迁移增强第一人称视频理解
- 在1.1M同步双视角数据上训练,显著提升任务表现
- 适合需要穿戴设备理解场景的智能助手研究者
AI个人助手通过机器人或可穿戴设备部署,需要具备具身理解能力以有效与人类协作。然而,现有多模态大语言模型(MLLM)主要关注第三人称(外视角)视觉,忽视了第一人称(内视角)视频的独特挑战。此外,数据获取成本高导致数据量小,制约了MLLM性能。为此,我们提出学习外视角与内视角之间的映射关系,利用现有MLLM中丰富的外视角知识来增强内视角视频理解。为此,我们构建了Ego-ExoClip预训练数据集,包含110万对来自Ego-Exo4D的同步内-外视角视频片段与文本对,并收集了多源指令微调数据集EgoIT以增强模型指令遵循能力。基于此,我们设计了三阶段渐进式映射学习流程:演示者自准备、演示者-学习者引导、学习者自练习。大量实验表明,现有MLLM在内视角视频理解上表现不佳,而我们的模型显著优于主流模型。
原文摘要 · Abstract (English)
AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric) vision, overlooking the unique challenges of first-person (egocentric) videos. Additionally, high acquisition costs limit data size, impairing MLLM performance. To address these challenges, we propose learning the mapping between exocentric and egocentric domains, leveraging the extensive exocentric knowledge within existing MLLMs to enhance egocentric video understanding. To this end, we introduce Ego-ExoClip, a pre-training dataset comprising 1.1M synchronized ego-exo clip-text pairs derived from Ego-Exo4D, together with the instruction-tuning dataset EgoIT, which is collected from multiple sources to enhance the model's instruction-following capabilities. Building upon the datasets, we propose a migration strategy and further design a progressive mapping learning pipeline with three stages: Demonstrator Self-Preparation, Demonstrator-Learner Guidance, and Learner Self-Practice. Extensive experiments across diverse egocentric tasks reveal that existing MLLMs perform inadequately in egocentric video understanding, while our model significantly outperforms these leading models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。