arXiv:2506.05287cs.CV2025-06NeurIPS被引 18

评测大模型在第一视角动态场景中识别、回忆和预测物体的能力。

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?

  • 构建包含3277个问答对的动态第一视角评测基准
  • 覆盖过去、现在、未来三类时间维度,评估11个细粒度维度
  • 适合研究具身智能、多模态大模型认知能力的学者使用

多模态大语言模型(MLLMs)的兴起推动了第一视角视觉应用的发展。这些应用需要对物体进行持续、上下文感知的理解,因为用户在动态且杂乱的环境中与工具互动。然而,现有具身评测基准主要关注静态场景探索,强调物体外观和空间属性,忽视了用户交互引发的动态变化评估。为弥补这一空白,我们提出EOC-Bench,一个创新的基准,用于系统评估动态第一视角场景中的以物体为中心的具身认知能力。EOC-Bench包含3,277个精心标注的问答对,分为过去、现在、未来三类时间类别,涵盖11个细粒度评估维度和3种视觉物体指代类型。为确保全面评估,我们设计了四种问题类型的混合格式人机协同标注框架,并提出一种新颖的多尺度时间精度指标,用于开放式时间评估。基于EOC-Bench,我们对多种专有、开源及基于物体的MLLM进行了综合评估。EOC-Bench成为提升MLLM具身物体认知能力的关键工具,为开发可靠的具身系统核心模型奠定了坚实基础。

原文摘要 · Abstract (English)

The emergence of multimodal large language models (MLLMs) has driven breakthroughs in egocentric vision applications. These applications necessitate persistent, context-aware understanding of objects, as users interact with tools in dynamic and cluttered environments. However, existing embodied benchmarks primarily focus on static scene exploration, emphasizing object's appearance and spatial attributes while neglecting the assessment of dynamic changes arising from users' interactions. To address this gap, we introduce EOC-Bench, an innovative benchmark designed to systematically evaluate object-centric embodied cognition in dynamic egocentric scenarios. Specially, EOC-Bench features 3,277 meticulously annotated QA pairs categorized into three temporal categories: Past, Present, and Future, covering 11 fine-grained evaluation dimensions and 3 visual object referencing types. To ensure thorough assessment, we develop a mixed-format human-in-the-loop annotation framework with four types of questions and design a novel multi-scale temporal accuracy metric for open-ended temporal evaluation. Based on EOC-Bench, we conduct comprehensive evaluations of various proprietary, open-source, and object-level MLLMs. EOC-Bench serves as a crucial tool for advancing the embodied object cognitive capabilities of MLLMs, establishing a robust foundation for developing reliable core models for embodied systems.

具身智能多模态模型动态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。