arXiv:2608.30983cs.RO2026-08

用自然语言描述任务,自动生成多样机械臂操作技能库。

Autonomously Acquiring Robot Manipulation Skills with Language-Driven Quality-Diversity

论文配图:Autonomously Acquiring Robot Manipulation Skills with Language-Driven Quality-Diversity
图 1 · 摘自论文原文
  • 用大模型自动设计评价与多样性指标,无需人工干预。
  • 在4个机器人操作任务中,生成的技能库优于传统方法。
  • 适合希望实现零样本适应的自主机器人研发者。

质量-多样性(QD)算法在机器人学习中日益受到关注,其生成的多样化运动基元库可使机器人在部署时实现零样本适应。然而,现有方法通常需专家设计成功条件、性能与多样性度量,严重限制了机器人的自主性。而现有的基于大模型的奖励塑造技术虽能实现自主学习,但仅输出单一高性能解,削弱了适应能力。本文提出一种新方法,仅需用自然语言描述任务,即可自主利用QD算法生成多样化的运动基元档案。为解决评价与多样性度量设计难题,我们提出一种自主探索机制,能可靠地生成覆盖性能与行为描述符(BD)空间的函数集。将策略探索建模为函数设计问题,降低维度后通过大模型采样,无需任务特定提示或微调。采用多行为描述符的MAP-Elites成功算法,有效利用异构的BD样本。在Genesis仿真器上的实验表明,该方法在4个机器人操作任务中均生成了高质量的多样化基元库,优于使用推断和手工参数化的经典QD算法。

原文摘要 · Abstract (English)

Quality-diversity (QD) algorithms have been gaining traction in robot learning, where diverse motion primitive libraries allow robots to adapt zero-shot to constraints at deployment time. However, such methods typically require expert designers to write the success condition, fitness and diversity metrics, and this strongly limits the robot's autonomy. On the other hand, existing LLM-based reward-shaping techniques allow robots to learn autonomously but only output single high-performing solutions, limiting the robot's adaptability. In this paper, we propose an approach designed to output diverse motion primitive archives by autonomously leveraging quality-diversity algorithms, only requiring a free-form description of the task in common language. To address the difficulty of designing relevant fitness and diversity metrics, we propose an autonomous exploration mechanism able to reliably output sets of functionals covering the fitness and behavior descriptor (BD) space. First, we pose policy exploration as a functional design problem, where the functional spaces are lower-dimensional than the full BD and fitness spaces, and propose an LLM-based exploration scheme to sample from these low-dimensional spaces without any task-specific prompts, fine-tuning or expert intervention. We adapt a multi-BD variant of the MAP-Elites success (MES) algorithm, designed to leverage the heterogeneous BD samples. Finally, through experiments based on the genesis simulator, we show that our method effectively generates archives of diverse motion primitives, outperforming classical QD algorithms with inferred and hand-written parametrizations on a set of $4$ robotic manipulation tasks.

机器人学习质量多样性大模型应用自主控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。