arXiv:2410.11623cs.CVcs.AI2024-10被引 34

构建首个第一人称视频理解评测基准,揭示大模型在真实场景中的能力短板。

VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI

  • 基于Ego4D数据集与GPT-4o自动生成四类任务数据,降低人工标注成本。
  • 三类模型在所有任务上表现均差,连GPT-4o也难以胜任第一人称视频理解。
  • 适合研究具身智能、多模态大模型与第一人称视觉的学者参考。

多模态大语言模型(MLLMs)的发展为具身智能应用开辟了新路径。在先前工作EgoThink基础上,本文提出VidEgoThink,一个全面评估第一人称视频理解能力的基准。为弥合MLLMs与具身智能低层控制之间的差距,设计了四个相互关联的任务:视频问答、层级规划、视觉定位和奖励建模。为减少人工标注成本,基于Ego4D数据集,利用GPT-4o的先验知识与多模态能力开发自动数据生成流水线。随后由三位人工标注者筛选数据以保证多样性与质量,最终形成VidEgoThink基准。对三类模型(API型MLLMs、开源图像型MLLMs、开源视频型MLLMs)进行广泛实验,结果表明,所有模型,包括GPT-4o,在各项第一人称视频理解任务中表现均不佳。这些发现表明,基础模型仍需重大改进才能有效应用于具身智能的第一人称场景。结论指出,VidEgoThink反映了将MLLMs用于第一人称视觉的研究趋势,旨在实现类人级主动观察与交互,适应复杂真实环境。

原文摘要 · Abstract (English)

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evaluating egocentric video understanding capabilities. To bridge the gap between MLLMs and low-level control in Embodied AI, we design four key interrelated tasks: video question-answering, hierarchy planning, visual grounding and reward modeling. To minimize manual annotation costs, we develop an automatic data generation pipeline based on the Ego4D dataset, leveraging the prior knowledge and multimodal capabilities of GPT-4o. Three human annotators then filter the generated data to ensure diversity and quality, resulting in the VidEgoThink benchmark. We conduct extensive experiments with three types of models: API-based MLLMs, open-source image-based MLLMs, and open-source video-based MLLMs. Experimental results indicate that all MLLMs, including GPT-4o, perform poorly across all tasks related to egocentric video understanding. These findings suggest that foundation models still require significant advancements to be effectively applied to first-person scenarios in Embodied AI. In conclusion, VidEgoThink reflects a research trend towards employing MLLMs for egocentric vision, akin to human capabilities, enabling active observation and interaction in the complex real-world environments.

具身智能第一人称视频多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。