arXiv:2606.02194cs.LG2026-06

用专家数据学奖励函数,让大模型在稀疏奖励下更高效优化

Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards

论文配图:Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards
图 1 · 摘自论文原文
  • 从专家演示中学习密集奖励,提升强化学习效率
  • 六项复杂操作任务中五项成功率超90%,优于稀疏奖励基线
  • 避免强化学习初期性能下降,适合机器人精细操作微调

使用行为克隆将专家示范数据蒸馏到大型生成模型中,是学习机器人控制能力的可扩展方法,尤其适用于灵巧操作。强化学习(RL)可用于利用额外经验进一步微调这些策略。一个未解问题是:RL是否比收集更多人类示范更具样本效率?先前工作通过在较小残差策略上应用RL,以可扩展方式微调预训练策略。然而,在典型稀疏奖励任务中,RL算法难以实现高效优化。我们探索逆强化学习,从专家示范中学习密集奖励函数,可能降低微调难度。具体考虑一致模仿学习,一种具备理论保证的逆强化学习方法,通过特定奖励形式改进行为克隆策略。我们证明该方法在六个稀疏操作任务中均保持或提升性能,且在五个复杂操作任务中成功率≥90%,优于使用稀疏奖励的基于RL的基线。通过确保初始预训练微调策略对初始奖励和评判器最优,我们的方法规避了强化学习微调中常见的性能下降,实现更快改进。

原文摘要 · Abstract (English)

Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation. Reinforcement learning (RL) can be used as a means to finetune these policies further using additional experience. An open question is whether RL is more sample-efficient than collecting more human demonstrations. Prior work has finetuned large pretrained policies in a scalable fashion by applying RL to a smaller residual policy that corrects the pretrained model. However, for the typical sparse reward tasks, RL algorithms can struggle to optimize the behavior in a sample-efficient manner. We explore inverse reinforcement learning, where a dense reward function is learned from expert demonstrations, potentially reducing the challenge of RL finetuning. We specifically consider coherent imitation learning, an IRL method that facilitates improvement of the BC policy through using a specific reward formulation with theoretical guarantees. We show that our IRL method maintains or improves the performance of pi-0.5 on all six sparse manipulation tasks and achieves a $\geq 90\%$ success rate on five out of six complex manipulation tasks, outperforming RL-based baselines using sparse rewards. By ensuring our initial pretrained finetuning policy is optimal for our initial reward and critic, our method circumvents the initial drop commonly seen in RL finetuning and enables faster improvement.

强化学习机器人控制逆强化学习稀疏奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。