arXiv:2412.12894cs.ROcs.LG2024-12被引 6

提出可计算均值的受限归一化流,提升强化学习策略的效率与可靠性。

Design of Restricted Normalizing Flow towards Arbitrary Stochastic Policy with Computational Efficiency

  • 设计受限归一化流(RNF),使变换可解析求均值
  • 用双峰t分布作基底,恢复表达能力,性能超越以往模型
  • 实机实验验证适用于实时机器人控制

本文提出一种基于归一化流(NF)的新式随机控制策略设计方法。在强化学习中,策略常被建模为可训练参数的概率分布;若表达能力不足,难以获得最优策略。混合模型虽具通用逼近能力,但冗余过多导致计算成本高,难以用于实时机器人控制。而归一化流虽具高表达性且计算开销较低,却因逆变换复杂无法解析计算均值,部署后仍保留随机行为,影响可靠性。为此,本文设计受限归一化流(RNF),通过限制逆变换结构实现均值的解析计算。同时,采用双峰t分布作为基底,弥补表达力损失,形成Bit-RNF。在强化学习基准测试中,Bit-RNF策略性能优于先前模型;实机实验进一步验证了其在真实场景中的适用性。相关视频已上传至YouTube:https://youtu.be/R_GJVZDW9bk。

原文摘要 · Abstract (English)

This paper proposes a new design method for a stochastic control policy using a normalizing flow (NF). In reinforcement learning (RL), the policy is usually modeled as a distribution model with trainable parameters. When this parameterization has less expressiveness, it would fail to acquiring the optimal policy. A mixture model has capability of a universal approximation, but it with too much redundancy increases the computational cost, which can become a bottleneck when considering the use of real-time robot control. As another approach, NF, which is with additional parameters for invertible transformation from a simple stochastic model as a base, is expected to exert high expressiveness and lower computational cost. However, NF cannot compute its mean analytically due to complexity of the invertible transformation, and it lacks reliability because it retains stochastic behaviors after deployment for robot controller. This paper therefore designs a restricted NF (RNF) that achieves an analytic mean by appropriately restricting the invertible transformation. In addition, the expressiveness impaired by this restriction is regained using bimodal student-t distribution as its base, so-called Bit-RNF. In RL benchmarks, Bit-RNF policy outperformed the previous models. Finally, a real robot experiment demonstrated the applicability of Bit-RNF policy to real world. The attached video is uploaded on youtube: https://youtu.be/R_GJVZDW9bk

强化学习归一化流策略设计机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。