arXiv:2605.11387cs.LGcs.RO2026-05被引 1

让机器人生成动作时既高效又多样,避免只学一种行为。

Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies

论文配图:Behavioral Mode Discovery for Fine-tuning Multimodal Generative Policies
图 1 · 摘自论文原文
  • 用无监督方法发现动作中的潜在行为模式。
  • 在强化学习中引入内在奖励,提升成功率并保持动作多样性。
  • 适合需要多策略的机器人任务,如复杂抓取与操作。

我们解决预训练生成式策略在强化学习微调过程中丧失动作分布多模性的难题。现有方法(如扩散策略)虽能提升任务表现,但常导致多种行为坍缩为单一高奖励模式。为此,我们提出一种无监督模式发现框架,挖掘生成式策略中的潜在行为模式。利用互信息作为内在奖励,正则化强化学习微调过程,在提升任务成功率的同时保持行为多样性。在机器人操作任务上的实验表明,本方法始终优于传统微调方法,取得更高成功率,并保留更丰富的多模态动作分布。

原文摘要 · Abstract (English)

We address the problem of fine-tuning pre-trained generative policies with reinforcement learning (RL) while preserving the multimodality of their action distributions. Existing methods for RL fine-tuning of generative policies (e.g., diffusion policies) improve task performance but often collapse diverse behaviors into a single reward-maximizing mode. To mitigate this issue, we propose an unsupervised mode discovery framework that uncovers latent behavioral modes within generative policies. The discovered modes enable the use of mutual information as an intrinsic reward, regularizing RL fine-tuning to enhance task success while maintaining behavioral diversity. Experiments on robotic manipulation tasks demonstrate that our method consistently outperforms conventional fine-tuning approaches, achieving higher success rates and preserving richer multimodal action distributions.

机器人控制生成策略多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。