提出统一生成框架UEPO,提升机器人离线到在线强化学习的泛化与鲁棒性
Opinion: Towards Unified Expressive Policy Optimization for Robust Robot Learning
- 基于多种子扩散策略,一次训练覆盖多种行为模式
- 动态差异正则化使策略多样性符合物理规律,提升适应能力
- 适合需要高鲁棒性的复杂机器人任务,如灵巧操作与运动控制
离线到在线强化学习(O2O-RL)为安全高效的机器人策略部署提供了新范式,但仍面临两大挑战:多模态行为覆盖不足,以及在线适应过程中的分布偏移。我们提出UEPO,一个受大语言模型预训练与微调策略启发的统一生成框架。贡献包括:(1) 多种子动态感知扩散策略,无需训练多个模型即可高效捕捉多样化行为;(2) 动态分歧正则化机制,强制策略多样性符合物理意义;(3) 基于扩散的数据增强模块,提升动态模型泛化能力。在D4RL基准上,UEPO在运动类任务上相较Uni-O4提升5.9%绝对性能,在灵巧操作任务上提升12.4%,展现出强泛化性与可扩展性。
原文摘要 · Abstract (English)
Offline-to-online reinforcement learning (O2O-RL) has emerged as a promising paradigm for safe and efficient robotic policy deployment but suffers from two fundamental challenges: limited coverage of multimodal behaviors and distributional shifts during online adaptation. We propose UEPO, a unified generative framework inspired by large language model pretraining and fine-tuning strategies. Our contributions are threefold: (1) a multi-seed dynamics-aware diffusion policy that efficiently captures diverse modalities without training multiple models; (2) a dynamic divergence regularization mechanism that enforces physically meaningful policy diversity; and (3) a diffusion-based data augmentation module that enhances dynamics model generalization. On the D4RL benchmark, UEPO achieves +5.9\% absolute improvement over Uni-O4 on locomotion tasks and +12.4\% on dexterous manipulation, demonstrating strong generalization and scalability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。