arXiv:2504.18160cs.LGcs.AI2025-04被引 2

让机器人模仿多样行为并可精准控制,告别单一输出。

Offline Learning of Controllable Diverse Behaviors

  • 构建时序一致的策略,确保全程行为连贯。
  • 在潜空间中实现行为可控,用户可选特定动作模式。
  • 支持多样化且可调控的智能体行为生成,适合复杂任务设计。

模仿学习(IL)旨在复现人类在特定任务中的行为。尽管传统方法在专家数据上表现高效,但通常只生成单一最优策略。近期研究尝试从多样化行为数据中学习,主要聚焦于过渡层面的多样性或轨迹级别的熵最大化,但这些方法难以充分还原演示数据的真实多样性,也无法实现可控的轨迹生成。为此,本文提出新方法,核心包含两个关键特性:a)时序一致性,确保整个轨迹中行为一致,而非仅在局部过渡处;b)可控制性,通过构建行为潜空间,使用户可根据需求选择激活特定行为。我们在多个任务与环境中对比了该方法与最先进方法的表现。项目页面:https://mathieu-petitbois.github.io/projects/swr/

原文摘要 · Abstract (English)

Imitation Learning (IL) techniques aim to replicate human behaviors in specific tasks. While IL has gained prominence due to its effectiveness and efficiency, traditional methods often focus on datasets collected from experts to produce a single efficient policy. Recently, extensions have been proposed to handle datasets of diverse behaviors by mainly focusing on learning transition-level diverse policies or on performing entropy maximization at the trajectory level. While these methods may lead to diverse behaviors, they may not be sufficient to reproduce the actual diversity of demonstrations or to allow controlled trajectory generation. To overcome these drawbacks, we propose a different method based on two key features: a) Temporal Consistency that ensures consistent behaviors across entire episodes and not just at the transition level as well as b) Controllability obtained by constructing a latent space of behaviors that allows users to selectively activate specific behaviors based on their requirements. We compare our approach to state-of-the-art methods over a diverse set of tasks and environments. Project page: https://mathieu-petitbois.github.io/projects/swr/

模仿学习行为控制多样性生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。