arXiv:2502.04692cs.ROcs.LG2025-02被引 4

用AI自动设计奖励函数,让机器人跑得更快更稳

STRIDE: Automating Reward Design, Deep Reinforcement Learning Training and Feedback Optimization in Humanoid Robotics Locomotion

  • 用大模型自动生成、评估并优化奖励函数,无需人工干预
  • 在复杂地形上实现冲刺级行走,效率比顶尖方法提升250%
  • 适合想快速迭代机器人控制算法的研究者和工程师

人形机器人在人工智能领域面临巨大挑战,需精确协调高自由度系统。深度强化学习(DRL)中的奖励函数设计仍是关键瓶颈,依赖大量人工、领域知识和反复调试。为此,我们提出STRIDE框架,基于智能体工程自动化完成人形机器人行走任务的奖励设计、DRL训练与反馈优化。结合大语言模型(LLMs)的代码生成、零样本生成与上下文优化能力,STRIDE无需任务特定提示或模板,即可生成、评估并迭代优化奖励函数。在多种包含不同人形机器人形态的环境中,STRIDE相较最先进框架EUREKA,在效率和任务表现上平均提升约250%。使用STRIDE生成的奖励,仿真人形机器人可在复杂地形实现冲刺级行走,彰显其对提升DRL工作流与人形机器人研究的推动作用。

原文摘要 · Abstract (English)

Humanoid robotics presents significant challenges in artificial intelligence, requiring precise coordination and control of high-degree-of-freedom systems. Designing effective reward functions for deep reinforcement learning (DRL) in this domain remains a critical bottleneck, demanding extensive manual effort, domain expertise, and iterative refinement. To overcome these challenges, we introduce STRIDE, a novel framework built on agentic engineering to automate reward design, DRL training, and feedback optimization for humanoid robot locomotion tasks. By combining the structured principles of agentic engineering with large language models (LLMs) for code-writing, zero-shot generation, and in-context optimization, STRIDE generates, evaluates, and iteratively refines reward functions without relying on task-specific prompts or templates. Across diverse environments featuring humanoid robot morphologies, STRIDE outperforms the state-of-the-art reward design framework EUREKA, achieving an average improvement of round 250% in efficiency and task performance. Using STRIDE-generated rewards, simulated humanoid robots achieve sprint-level locomotion across complex terrains, highlighting its ability to advance DRL workflows and humanoid robotics research.

强化学习机器人控制自动奖励设计大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。