arXiv:2409.17523cs.CVcs.AI2024-09中稿 · ACMMM 24被引 18

构建首个大规模第一人称视频多任务理解框架,统一动作识别与流程学习。

EAGLE: Egocentric AGgregated Language-video Engine

  • 基于40万样本构建首个第一人称视频指令微调数据集
  • 多模态大模型在400K数据上实现跨任务统一理解
  • 适合研究第一人称视频与多模态大模型的学者

第一人称视频分析的快速发展为从第一视角理解人类行为与意图提供了新洞察。尽管取得进展,动作识别、流程学习和时刻检索等任务的碎片化,以及标注不一致和模型孤立发展,阻碍了对视频内容的整体理解。为此,我们提出EAGLE(第一人称聚合语言-视频引擎)模型和EAGLE-400K数据集,构建一个整合多种第一人称视频理解任务的统一框架。EAGLE-400K是首个专为第一人称视频设计的大规模指令微调数据集,包含400,000个多样化样本,可提升从动作识别到流程知识学习的广泛任务表现。此外,EAGLE是一种强大的视频多模态大语言模型(MLLM),能有效捕捉空间与时间信息。我们还提出一套评估指标,以全面评估MLLM在第一人称视频理解中的性能。大量实验表明,EAGLE在现有模型中表现更优,展现出在任务特定理解与整体视频解释间的平衡能力。通过EAGLE,我们旨在为真实场景中的研究机遇与应用铺平道路。

原文摘要 · Abstract (English)

The rapid evolution of egocentric video analysis brings new insights into understanding human activities and intentions from a first-person perspective. Despite this progress, the fragmentation in tasks like action recognition, procedure learning, and moment retrieval, \etc, coupled with inconsistent annotations and isolated model development, hinders a holistic interpretation of video content. In response, we introduce the EAGLE (Egocentric AGgregated Language-video Engine) model and the EAGLE-400K dataset to provide a unified framework that integrates various egocentric video understanding tasks. EAGLE-400K, the \textit{first} large-scale instruction-tuning dataset tailored for egocentric video, features 400K diverse samples to enhance a broad spectrum of tasks from activity recognition to procedure knowledge learning. Moreover, EAGLE, a strong video multimodal large language model (MLLM), is designed to effectively capture both spatial and temporal information. In addition, we propose a set of evaluation metrics designed to facilitate a thorough assessment of MLLM for egocentric video understanding. Our extensive experiments demonstrate EAGLE's superior performance over existing models, highlighting its ability to balance task-specific understanding with holistic video interpretation. With EAGLE, we aim to pave the way for research opportunities and practical applications in real-world scenarios.

第一人称视频多模态大模型任务统一指令微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。