arXiv:2507.06157cs.ROcs.CL2025-07

用大模型评估机器人协作任务,发现推理模型表现更优。

Evaluation of Habitat Robotics using Large Language Models

  • 用大语言模型在随机厨房场景中测试双机器人协作任务
  • 推理型模型o3-mini在各类配置下均优于GPT-4o和Llama 3
  • 适合关注具身智能与多智能体协作的研究者

本文聚焦于使用Meta PARTNER基准评估大语言模型在具身机器人任务中的表现。PARTNER提供简化的环境与随机化室内厨房场景中的机器人交互。每个场景分配一项任务,由两个机器人协同完成。我们在多个前沿模型上进行了评估。结果表明,在具备中央控制、分布式控制、全可观测性及部分可观测性配置下,推理型模型OpenAI o3-mini的表现均优于非推理型模型如GPT-4o和Llama 3。该研究为具身机器人发展提供了有前景的新方向。

原文摘要 · Abstract (English)

This paper focuses on evaluating the effectiveness of Large Language Models at solving embodied robotic tasks using the Meta PARTNER benchmark. Meta PARTNR provides simplified environments and robotic interactions within randomized indoor kitchen scenes. Each randomized kitchen scene is given a task where two robotic agents cooperatively work together to solve the task. We evaluated multiple frontier models on Meta PARTNER environments. Our results indicate that reasoning models like OpenAI o3-mini outperform non-reasoning models like OpenAI GPT-4o and Llama 3 when operating in PARTNR's robotic embodied environments. o3-mini displayed outperform across centralized, decentralized, full observability, and partial observability configurations. This provides a promising avenue of research for embodied robotic development.

具身智能多智能体大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。