arXiv:2510.21860cs.ROcs.AI2025-10被引 1

测试大模型控制机器人应对真实世界混乱的能力,发现人类仍远超模型。

Butter-Bench: Evaluating LLM Controlled Robots for Practical Intelligence

  • 单独评估大模型的高层决策能力,剥离低层控制影响。
  • 顶尖模型仅达40%正确率,人类平均95%。
  • 模型在多步空间规划与社交理解上表现最差,微调无效。

我们提出Butter-Bench,一个用于评估大语言模型(LLM)控制机器人在现实物理世界中实用智能的基准。实用智能指应对真实世界混乱的能力。当前先进机器人系统采用分层架构:大模型负责高层推理,视觉语言动作(VLA)模型负责底层控制。Butter-Bench独立评估大模型部分。尽管大模型在需要分析智能的任务中多次超越人类,但在Butter-Bench上,人类仍显著领先。最佳大模型得分40%,而人类平均得分为95%。大模型在多步空间规划和社交理解方面表现最弱。我们还评估了针对具身推理微调的大模型,发现其在Butter-Bench上的表现未获提升。

原文摘要 · Abstract (English)

We present Butter-Bench, a benchmark evaluating large language model (LLM) controlled robots for practical intelligence, defined as the ability to navigate the messiness of the physical world. Current state-of-the-art robotic systems use a hierarchical architecture with LLMs in charge of high-level reasoning, and a Vision Language Action (VLA) model for low-level control. Butter-Bench evaluates the LLM part in isolation from the VLA. Although LLMs have repeatedly surpassed humans in evaluations requiring analytical intelligence, we find humans still outperform LLMs on Butter-Bench. The best LLMs score 40% on Butter-Bench, while the mean human score is 95%. LLMs struggled the most with multi-step spatial planning and social understanding. We also evaluate LLMs that are fine-tuned for embodied reasoning and conclude that this training does not improve their score on Butter-Bench.

机器人大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。