用局部传输映射优化流策略,解决离线强化学习中的对齐偏差问题。
Fisher Decorator: Refining Flow Policy via a Local Transport Map

- 引入局部运输映射,以残差位移修正初始流策略
- 基于Fisher信息矩阵实现可控制的近似误差,优于原有方法
- 适合追求高精度策略优化的离线强化学习研究者
基于流的离线强化学习近期通过流匹配参数化策略取得了优异性能,但仍面临表达能力、最优性与效率之间的权衡。现有方法将$L_2$正则视为$W_2$距离的上界,但在离线设置下存在根本性几何错配:行为策略流形具有各向异性,而$L_2$或$W_2$上界正则为各向同性且对密度不敏感,导致优化方向系统性偏离。本文从几何视角重新审视离线强化学习,提出将策略精炼建模为局部运输映射——在初始流策略基础上添加残差位移。通过分析诱导的密度变换,推导出由Fisher信息矩阵主导的KL约束目标的局部二次逼近,从而获得可解析的各向异性优化形式。利用流速度中嵌入的得分函数,得到高效优化对应的二次约束。结果表明,先前方法的最优性差距源于其各向同性近似。相比之下,本框架在最优解邻域内具备可控的近似误差。大量实验验证了其在多种离线强化学习基准上的领先性能。
原文摘要 · Abstract (English)
Recent advances in flow-based offline reinforcement learning (RL) have achieved strong performance by parameterizing policies via flow matching. However, they still face critical trade-offs among expressiveness, optimality, and efficiency. In particular, existing flow policies interpret the $L_2$ regularization as an upper bound of the 2-Wasserstein distance ($W_2$), which can be problematic in offline settings. This issue stems from a fundamental geometric mismatch: the behavioral policy manifold is inherently anisotropic, whereas the $L_2$ (or upper bound of $W_2$) regularization is isotropic and density-insensitive, leading to systematically misaligned optimization directions. To address this, we revisit offline RL from a geometric perspective and show that policy refinement can be formulated as a local transport map: an initial flow policy augmented by a residual displacement. By analyzing the induced density transformation, we derive a local quadratic approximation of the KL-constrained objective governed by the Fisher information matrix, enabling a tractable anisotropic optimization formulation. By leveraging the score function embedded in the flow velocity, we obtain a corresponding quadratic constraint for efficient optimization. Our results reveal that the optimality gap in prior methods arises from their isotropic approximation. In contrast, our framework achieves a controllable approximation error within a provable neighborhood of the optimal solution. Extensive experiments demonstrate state-of-the-art performance across diverse offline RL benchmarks. See project page: https://github.com/ARC0127/Fisher-Decorator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。