arXiv:2409.18127cs.CV2024-09CVPR被引 30

用多模态数据统一理解第一视角动作,让机器像人一样看懂日常动作。

EgoLM: Multi-Modal Language Model of Egocentric Motions

  • 基于大语言模型建模动作与语言的联合分布,融合视频和传感器数据
  • 在大规模多模态数据集上实现动作生成与描述的高精度统一建模
  • 适合可穿戴设备、智能助手等需要理解第一视角行为的场景

随着可穿戴设备的普及,学习第一视角动作对发展上下文感知的人工智能至关重要。本文提出 EgoLM,一个从多模态输入(如第一视角视频和运动传感器)中追踪与理解第一视角动作的通用框架。EgoLM 利用丰富的上下文信息来消除单模态条件下动作理解的歧义性。核心思想是利用大语言模型(LLM)建模第一视角动作与自然语言的联合分布,将多模态传感器输入编码并投影至语言模型的联合隐空间,用于提示动作生成或文本生成,分别实现动作追踪与理解。在大规模多模态人体动作数据集上的大量实验验证了 EgoLM 作为通用第一视角学习模型的有效性。

原文摘要 · Abstract (English)

As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and understands egocentric motions from multi-modal inputs, e.g., egocentric videos and motion sensors. EgoLM exploits rich contexts for the disambiguation of egomotion tracking and understanding, which are ill-posed under single modality conditions. To facilitate the versatile and multi-modal framework, our key insight is to model the joint distribution of egocentric motions and natural languages using large language models (LLM). Multi-modal sensor inputs are encoded and projected to the joint latent space of language models, and used to prompt motion generation or text generation for egomotion tracking or understanding, respectively. Extensive experiments on large-scale multi-modal human motion dataset validate the effectiveness of EgoLM as a generalist model for universal egocentric learning.

第一视角多模态动作理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。