arXiv:2502.03717cs.ROcs.AI2025-02ICRA被引 3

用语言引导偏好学习,四次提问就让四足机器人做出自然动作

Efficiently Generating Expressive Quadruped Behaviors via Language-Guided Preference Learning

  • 用预训练大模型生成初始动作样本,再通过人类偏好反馈优化
  • 仅需4次交互即可学会精准表达的动作行为,效率远超传统方法
  • 适合希望快速生成自然人机互动机器人的研究者和开发者

富有表现力的机器人行为对机器人在社交环境中的普及至关重要。近期基于学习的四足行走控制器已实现更动态、多样化的动作。然而,在不同场景下针对不同用户确定最优行为仍是挑战。现有方法要么依赖自然语言输入(高效但分辨率低),要么通过人类偏好学习(分辨率高但样本效率低)。本文提出一种新方法——语言引导偏好学习(LGPL),利用预训练大模型生成初始行为样本,并通过偏好反馈进行精炼,使行为更贴近人类预期。核心洞察是:大模型可引导偏好学习的采样过程,显著提升样本效率。实验表明,LGPL仅需4次查询即可快速学习准确且富有表现力的行为,优于纯语言参数化模型和传统偏好学习方法。

原文摘要 · Abstract (English)

Expressive robotic behavior is essential for the widespread acceptance of robots in social environments. Recent advancements in learned legged locomotion controllers have enabled more dynamic and versatile robot behaviors. However, determining the optimal behavior for interactions with different users across varied scenarios remains a challenge. Current methods either rely on natural language input, which is efficient but low-resolution, or learn from human preferences, which, although high-resolution, is sample inefficient. This paper introduces a novel approach that leverages priors generated by pre-trained LLMs alongside the precision of preference learning. Our method, termed Language-Guided Preference Learning (LGPL), uses LLMs to generate initial behavior samples, which are then refined through preference-based feedback to learn behaviors that closely align with human expectations. Our core insight is that LLMs can guide the sampling process for preference learning, leading to a substantial improvement in sample efficiency. We demonstrate that LGPL can quickly learn accurate and expressive behaviors with as few as four queries, outperforming both purely language-parameterized models and traditional preference learning approaches. Website with videos: https://lgpl-gaits.github.io/

四足机器人语言控制偏好学习高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。