arXiv:2410.07177cs.CVcs.AI2024-10中稿 · ICLR被引 27

构建首个面向第一人称视频理解的多模态大模型,提升长视频细节记忆能力。

MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA

  • 自动生成700万条第一人称视频QA数据,覆盖30秒至1小时长视频。
  • 提出629段视频、7026个问题的新基准,支持长视频细节识别评测。
  • 创新记忆指针提示机制,增强对长视频全局与关键信息的理解。

本研究致力于构建用于第一人称视频理解的多模态基础模型。首先,针对第一人称视频问答数据稀缺的问题,基于人工标注数据自动生成了700万条高质量、时长在30秒至1小时之间的第一人称视频QA样本,构成目前规模最大的第一人称问答数据集。其次,构建了一个包含629段视频和7,026个问题的挑战性评测基准,用于评估模型在不同长度视频中识别与记忆视觉细节的能力,并引入去偏评估方法以缓解模型固有的语言偏差。第三,提出一种专用多模态架构,包含新颖的“记忆指针提示”机制:通过全局概览步骤获取视频整体理解并定位关键视觉信息,再利用该信息生成回答,从而更有效地理解长视频内容。基于这些数据、基准与模型,我们构建了MM-Ego,一个在第一人称视频理解任务上表现卓越的第一人称多模态大模型。

原文摘要 · Abstract (English)

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric video understanding, we automatically generate 7M high-quality QA samples for egocentric videos ranging from 30 seconds to one hour long in Ego4D based on human-annotated data. This is one of the largest egocentric QA datasets. Second, we contribute a challenging egocentric QA benchmark with 629 videos and 7,026 questions to evaluate the models' ability in recognizing and memorizing visual details across videos of varying lengths. We introduce a new de-biasing evaluation method to help mitigate the unavoidable language bias present in the models being evaluated. Third, we propose a specialized multimodal architecture featuring a novel "Memory Pointer Prompting" mechanism. This design includes a \textit{global glimpse} step to gain an overarching understanding of the entire video and identify key visual information, followed by a fallback step that utilizes the key visual information to generate responses. This enables the model to more effectively comprehend extended video content. With the data, benchmark, and model, we build MM-Ego, an egocentric multimodal LLM that shows powerful performance on egocentric video understanding.

第一人称视频多模态大模型视频问答长视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。