用流模型提升分层强化学习的数据效率,让机器人更少试错也能完成复杂任务。
Data-Efficient Hierarchical Goal-Conditioned Reinforcement Learning via Normalizing Flows
- 用流模型替代传统高斯策略,实现多模态行为建模与高效采样
- 在有限数据下优于已有分层方法,在多个长程任务中表现领先
- 适合研究数据稀缺场景下的智能体决策与分层强化学习
分层目标条件强化学习(H-GCRL)通过将复杂长程任务分解为结构化子目标,提供强大求解框架。但其实际应用受限于数据效率低和策略表达能力弱,尤其在离线或数据稀疏场景。本文提出基于流模型的分层隐式Q学习(NF-HIQL),在高层与低层均采用表达能力强的流模型策略,实现可计算的对数似然、高效采样,并能捕捉丰富多模态行为。理论分析给出真实非体积保持(RealNVP)策略的明确KL散度界与类PAC样本效率结果,证明该方法在保持稳定性的同时提升泛化能力。实验在OGBench的运动、控球及多步操作等多样化长程任务上验证,NF-HIQL持续优于现有目标条件与分层基线,展现出在数据有限时的更强鲁棒性,凸显流架构在可扩展、数据高效的分层强化学习中的潜力。
原文摘要 · Abstract (English)
Hierarchical goal-conditioned reinforcement learning (H-GCRL) provides a powerful framework for tackling complex, long-horizon tasks by decomposing them into structured subgoals. However, its practical adoption is hindered by poor data efficiency and limited policy expressivity, especially in offline or data-scarce regimes. In this work, Normalizing flow-based hierarchical implicit Q-learning (NF-HIQL), a novel framework that replaces unimodal gaussian policies with expressive normalizing flow policies at both the high- and low-levels of the hierarchy is introduced. This design enables tractable log-likelihood computation, efficient sampling, and the ability to model rich multimodal behaviors. New theoretical guarantees are derived, including explicit KL-divergence bounds for Real-valued non-volume preserving (RealNVP) policies and PAC-style sample efficiency results, showing that NF-HIQL preserves stability while improving generalization. Empirically, NF-HIQL is evaluted across diverse long-horizon tasks in locomotion, ball-dribbling, and multi-step manipulation from OGBench. NF-HIQL consistently outperforms prior goal-conditioned and hierarchical baselines, demonstrating superior robustness under limited data and highlighting the potential of flow-based architectures for scalable, data-efficient hierarchical reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。