arXiv:2601.21570cs.AIcs.RO2026-01被引 1

用大模型自动训练机器人,让智能体自己调优策略。

From Digital to Physical: Digital Agents as Autonomous Coaches for Physical Intelligence

  • 让大模型通过环境反馈反复试错,自主设计机器人控制策略。
  • 自动优化后成功率比人工设计高26.5%,接近闭源模型水平。
  • 能自我纠错,从几乎失败的状态中恢复任务性能。

具身智能领域正快速向通用机器人系统演进,依赖高保真仿真与大规模数据采集。然而,这一扩展能力严重受限于对人工干预的依赖,从奖励设计到超参数调优均需大量人力。受大语言模型在软件自动化与科学发现中的启发,我们提出 extsc{EmboCoach-Bench} 基准,评估大模型智能体自主构建具身策略的能力。该框架涵盖32个专家精选的强化学习与模仿学习任务,以可执行代码为统一接口。我们突破静态生成,评估动态闭环工作流:智能体利用环境反馈迭代撰写、调试与优化解决方案,涵盖物理感知奖励设计与扩散策略等架构改进。大量实验揭示三个关键洞察:(1) 自主智能体在平均成功率上较人工基线提升26.5%;(2) 带环境反馈的智能体工作流显著增强策略开发,大幅缩小开源与专有模型间的性能差距;(3) 智能体具备自纠正能力,可在病理工程场景下通过模拟闭环调试,从近乎完全失败中恢复任务表现。本研究为自演化具身智能奠定基础,推动具身智能从人工调参迈向可扩展的自主工程范式。

原文摘要 · Abstract (English)

The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection. However, this scaling capability remains severely bottlenecked by a reliance on labor-intensive manual oversight from intricate reward shaping to hyperparameter tuning across heterogeneous backends. Inspired by LLMs' success in software automation and science discovery, we introduce \textsc{EmboCoach-Bench}, a benchmark evaluating the capacity of LLM agents to autonomously engineer embodied policies. Spanning 32 expert-curated RL and IL tasks, our framework posits executable code as the universal interface. We move beyond static generation to assess a dynamic closed-loop workflow, where agents leverage environment feedback to iteratively draft, debug, and optimize solutions, spanning improvements from physics-informed reward design to policy architectures such as diffusion policies. Extensive evaluations yield three critical insights: (1) autonomous agents can qualitatively surpass human-engineered baselines by 26.5\% in average success rate; (2) agentic workflow with environment feedback effectively strengthens policy development and substantially narrows the performance gap between open-source and proprietary models; and (3) agents exhibit self-correction capabilities for pathological engineering cases, successfully resurrecting task performance from near-total failures through iterative simulation-in-the-loop debugging. Ultimately, this work establishes a foundation for self-evolving embodied intelligence, accelerating the paradigm shift from labor-intensive manual tuning to scalable, autonomous engineering in embodied AI field.

具身智能大模型代理自动调优

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。