arXiv:2602.11087cs.LGcs.AI2026-02中稿 · AAMAS 2026

提出可自适应约束的广义f散度,提升低探索数据下的离线强化学习性能

General Flexible $f$-divergence for Challenging Offline RL Datasets with Low Stochasticity and Diverse Behavior Policies

  • 基于线性规划与凸共轭,建立f散度与贝尔曼残差约束的联系
  • 在MuJoCo等环境上验证,新方法显著提升挑战性数据集上的学习效果
  • 适合处理多行为策略、探索不足的复杂离线数据集

离线强化学习旨在改进数据生成的行为策略,同时将学习策略限制在数据集支持范围内。但实际离线数据集常存在探索不足、行为策略多样且能力差异大等问题。探索不足会影响Q或V值估计,而对多种行为策略的严格约束又过于保守。为此,本文通过更一般的线性规划形式与凸共轭,揭示f散度与贝尔曼残差约束之间的关联,提出广义灵活f散度函数,使约束能根据数据集特性自适应调整。在MuJoCo、Fetch和AdroitHand环境上的实验表明,该方法在兼容的约束优化算法中有效提升了在挑战性数据集上的学习性能。

原文摘要 · Abstract (English)

Offline RL algorithms aim to improve upon the behavior policy that produces the collected data while constraining the learned policy to be within the support of the dataset. However, practical offline datasets often contain examples with little diversity or limited exploration of the environment, and from multiple behavior policies with diverse expertise levels. Limited exploration can impair the offline RL algorithm's ability to estimate \textit{Q} or \textit{V} values, while constraining towards diverse behavior policies can be overly conservative. Such datasets call for a balance between the RL objective and behavior policy constraints. We first identify the connection between $f$-divergence and optimization constraint on the Bellman residual through a more general Linear Programming form for RL and the convex conjugate. Following this, we introduce the general flexible function formulation for the $f$-divergence to incorporate an adaptive constraint on algorithms' learning objectives based on the offline training dataset. Results from experiments on the MuJoCo, Fetch, and AdroitHand environments show the correctness of the proposed LP form and the potential of the flexible $f$-divergence in improving performance for learning from a challenging dataset when applied to a compatible constrained optimization algorithm.

离线RLf散度策略约束强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。