arXiv:2508.06553cs.CV2025-08被引 2

用静态场景统一评估具身智能,降低门槛提升效率

Static and Plugged: Make Embodied Evaluation Simple

  • 用静态场景替代交互环境,实现简单插拔式评估
  • 覆盖42个场景、8个维度,评估19个VLM和11个VLA模型
  • 提供首个统一静态基准榜单,适合研究者快速对比模型

具身智能快速发展,亟需高效评估方法。现有基准多依赖交互式仿真环境或真实世界部署,成本高、碎片化且难扩展。为此,我们提出StaticEmbodiedBench,一个即插即用的基准,通过静态场景表示实现统一评估。涵盖42种多样化场景和8个核心维度,支持通过简易接口进行可扩展、全面的评估。我们评估了19个视觉-语言模型(VLMs)和11个视觉-语言-动作模型(VLAs),建立了首个具身智能的静态统一排行榜。此外,我们公开了200个样本子集,以加速具身智能发展。

原文摘要 · Abstract (English)

Embodied intelligence is advancing rapidly, driving the need for efficient evaluation. Current benchmarks typically rely on interactive simulated environments or real-world setups, which are costly, fragmented, and hard to scale. To address this, we introduce StaticEmbodiedBench, a plug-and-play benchmark that enables unified evaluation using static scene representations. Covering 42 diverse scenarios and 8 core dimensions, it supports scalable and comprehensive assessment through a simple interface. Furthermore, we evaluate 19 Vision-Language Models (VLMs) and 11 Vision-Language-Action models (VLAs), establishing the first unified static leaderboard for Embodied intelligence. Moreover, we release a subset of 200 samples from our benchmark to accelerate the development of embodied intelligence.

具身智能评估基准静态场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。