arXiv:2510.12072cs.AIcs.RO2025-10被引 3

用仿真环境训练大模型,让其学会真实世界中的决策能力。

EmboMatrix: A Scalable Training-Ground for Embodied Decision-Making

  • 构建仿真训练场,支持大规模任务与物理交互模拟。
  • 70亿参数模型在两个基准上比6710亿参数基线高9.5%。
  • 适合研究具身智能、大模型与物理世界交互的学者。

具身决策使智能体通过与物理世界的持续互动,将高层目标转化为可执行动作,是通用具身智能的核心。大语言模型(LLM)虽具备通用决策能力,但仅基于语言训练缺乏对物理环境的真实理解。为此,我们提出“训练场”概念——一个集任务与场景仿真、具身交互和反馈信号于一体的综合性基础设施,为LLM提供获取真实具身理解的一站式解决方案。本文提出首个此类训练场EmboMatrix,支持海量多样化任务、高效仿真与精确奖励。EmboMatrix采用多智能体数据引擎生成大规模任务与场景,分布式异构硬件系统实现可扩展仿真,并设计多层级奖励架构实现精准监督。基于EmboMatrix,我们训练出EmboBrain,一个通过大量具身交互涌现出决策能力的LLM。实验表明,EmboBrain-7B在两个挑战性具身决策基准上超越6710亿参数的DeepSeek-R1基线9.5%,验证了基于环境交互学习在构建真正智能具身代理中的强大作用。

原文摘要 · Abstract (English)

Embodied decision-making enables agents to translate high-level goals into executable actions through continuous interactions within the physical world, forming a cornerstone of general-purpose embodied intelligence. Large language models (LLMs), with their general decision-making capabilities, offer a promising path to realize this potential; however, LLMs trained solely on language lack exposure to physical environments, limiting their true embodied understanding. To bridge this gap, we propose the concept of a training ground: a comprehensive infrastructure that provides task and scene simulation, embodied interaction, and feedback signals, offering a one-stop solution for LLM acquire genuine embodied decision-making skills. In this work, we present EmboMatrix, the first training ground of its kind, providing massive and diverse tasks with efficient simulation and precise rewards. EmboMatrix incorporates a series of novel techniques: a multi-agent data engine for large-scale task and scene generation, a distributed heterogeneous-hardware system for scalable simulation, and a multi-level reward architecture for precise supervision. Leveraging EmboMatrix, we cultivate EmboBrain, an LLM whose embodied decision-making abilities emerge from extensive embodied interactions. Experiments show that EmboBrain-7B surpasses the 671B DeepSeek-R1 baseline by 9.5\% on two challenging embodied decision-making benchmarks, demonstrating the power of interactive, environment-grounded learning for building truly intelligent embodied agents.

具身智能大模型仿真训练决策生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。