arXiv:2501.05031cs.CVcs.LG2025-01CVPR被引 31

构建首个系统评估视觉语言模型具身认知能力的基准测试

ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark

  • 设计多模态具身认知评测框架,覆盖30个认知维度
  • 通过人工标注与多轮筛选确保数据高质量与强视觉依赖性
  • 适合研究具身智能、机器人认知与多模态模型评估的学者

大型视觉语言模型(LVLMs)在提升机器人泛化能力方面日益显著。因此,基于第一人称视频的LVLM具身认知能力备受关注。然而,当前具身视频问答数据集缺乏全面系统的评估框架,关键具身认知问题如机器人自我认知、动态场景感知和幻觉等很少被关注。为此,我们提出ECBench,一个高质量基准,用于系统评估LVLM的具身认知能力。ECBench涵盖多样化场景视频源、开放多样的问题形式,包含30个具身认知维度。为保障质量、平衡性与高视觉依赖性,采用类无关的人工精细标注及多轮问题筛选策略。此外,我们引入ECEval,一个综合评估系统,确保指标公平合理。利用ECBench,我们对专有、开源及任务特定的LVLM进行了广泛评估。ECBench对推动LVLM具身认知能力发展具有重要意义,为构建可靠具身智能核心模型奠定基础。所有数据与代码已公开于https://github.com/Rh-Dang/ECBench。

原文摘要 · Abstract (English)

The enhancement of generalization in robots by large vision-language models (LVLMs) is increasingly evident. Therefore, the embodied cognitive abilities of LVLMs based on egocentric videos are of great interest. However, current datasets for embodied video question answering lack comprehensive and systematic evaluation frameworks. Critical embodied cognitive issues, such as robotic self-cognition, dynamic scene perception, and hallucination, are rarely addressed. To tackle these challenges, we propose ECBench, a high-quality benchmark designed to systematically evaluate the embodied cognitive abilities of LVLMs. ECBench features a diverse range of scene video sources, open and varied question formats, and 30 dimensions of embodied cognition. To ensure quality, balance, and high visual dependence, ECBench uses class-independent meticulous human annotation and multi-round question screening strategies. Additionally, we introduce ECEval, a comprehensive evaluation system that ensures the fairness and rationality of the indicators. Utilizing ECBench, we conduct extensive evaluations of proprietary, open-source, and task-specific LVLMs. ECBench is pivotal in advancing the embodied cognitive capabilities of LVLMs, laying a solid foundation for developing reliable core models for embodied agents. All data and code are available at https://github.com/Rh-Dang/ECBench.

具身智能多模态模型评测基准视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。