arXiv:2606.25398cs.RO2026-06

用自然语言指导机器人行走,自动学习多目标奖励函数。

MAPL: Multi-Objective Preference Learning for Robot Locomotion

论文配图:MAPL: Multi-Objective Preference Learning for Robot Locomotion
图 1 · 摘自论文原文
  • 通过大模型对轨迹按语义维度独立打分,实现多目标偏好学习。
  • 在4个四足环境上性能媲美甚至超过人工设计奖励函数。
  • 无需领域知识,仅需自然语言描述即可训练高性能策略。

奖励设计仍是机器人行走强化学习中的主要瓶颈,高质量策略通常依赖精心调优的任务特定奖励函数。基于偏好的强化学习提供了一种替代方案,但现有基于大模型的方法通常仅要求对行为进行整体判断,难以捕捉高质量行走所涉及的多重竞争目标。我们提出多目标人工智能引导偏好学习(MAPL),该框架从高层次自然语言目标中学习行走奖励,而非手动构建奖励方程。MAPL通过大模型在语义有意义的维度上独立比较轨迹,使用与地形无关、几乎无需领域知识的通用语言描述。各目标偏好用于训练多头偏好评分模型,其输出聚合为策略优化的标量奖励。在四个四足行走环境中,MAPL仅使用大模型生成的偏好训练策略,性能达到或超过专家设计奖励,同时完全消除任务特定的奖励工程。

原文摘要 · Abstract (English)

Reward design remains a major bottleneck in reinforcement learning for robot locomotion, where successful policies often depend on carefully tuned, task-specific reward functions. Preference-based reinforcement learning offers an alternative, but existing LLM-based methods typically ask for a single overall judgment between behaviors, making it difficult to capture the multiple competing objectives that underlie high-quality locomotion. We present Multi-Objective AI-Informed Preference Learning (MAPL), a framework that learns locomotion rewards from high-level natural language objectives rather than manually engineered reward equations. MAPL prompts a large language model to compare trajectories independently along semantically meaningful criteria, using generic language descriptions that are terrain-invariant and require little domain expertise. These objective-wise preferences are used to train a multi-head preference scoring model, whose outputs are aggregated to form a scalar reward for policy optimization. Across four quadruped locomotion environments, MAPL trains policies using only LLM-generated preferences and achieves performance comparable to or better than expert-designed rewards, while eliminating task-specific reward engineering.

强化学习机器人控制大模型应用多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。