用大模型自动生成偏好标签,让机器人高效学习复杂行为。
LAPP: Large Language Model Feedback for Preference-Driven Reinforcement Learning
- 用大模型从轨迹中自动生成偏好标签,替代人工标注
- 在四足和灵巧操作任务中实现更快学习与更高性能
- 适合希望快速定制机器人行为的开发者与研究者
我们提出大型语言模型辅助的偏好预测(LAPP)框架,用于机器人学习。该框架通过大语言模型(LLM)自动从强化学习过程中收集的原始状态-动作轨迹中生成偏好标签,无需依赖复杂的奖励工程、人类示范、动作捕捉或昂贵的成对偏好标注。这些标签用于训练在线偏好预测器,进而引导策略优化以满足人类提供的高层次行为规范。核心贡献在于将大模型融入强化学习反馈环,实现基于轨迹层面的偏好预测,使机器人能掌握复杂技能,包括对步态模式和节奏时间的细微控制。我们在多种四足运动与灵巧操作任务上评估了LAPP,结果表明其具备高效学习能力、更高的最终性能、更快的适应速度以及对高层行为的精确控制。尤其值得注意的是,LAPP成功实现了四足后空翻等高度动态且富有表现力的任务,而这些任务在传统大模型生成或手工设计奖励下仍难以达成。结果表明LAPP是可扩展的偏好驱动机器人学习的重要方向。
原文摘要 · Abstract (English)
We introduce Large Language Model-Assisted Preference Prediction (LAPP), a novel framework for robot learning that enables efficient, customizable, and expressive behavior acquisition with minimum human effort. Unlike prior approaches that rely heavily on reward engineering, human demonstrations, motion capture, or expensive pairwise preference labels, LAPP leverages large language models (LLMs) to automatically generate preference labels from raw state-action trajectories collected during reinforcement learning (RL). These labels are used to train an online preference predictor, which in turn guides the policy optimization process toward satisfying high-level behavioral specifications provided by humans. Our key technical contribution is the integration of LLMs into the RL feedback loop through trajectory-level preference prediction, enabling robots to acquire complex skills including subtle control over gait patterns and rhythmic timing. We evaluate LAPP on a diverse set of quadruped locomotion and dexterous manipulation tasks and show that it achieves efficient learning, higher final performance, faster adaptation, and precise control of high-level behaviors. Notably, LAPP enables robots to master highly dynamic and expressive tasks such as quadruped backflips, which remain out of reach for standard LLM-generated or handcrafted rewards. Our results highlight LAPP as a promising direction for scalable preference-driven robot learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。