arXiv:2602.19917cs.LGcs.RO2026-02被引 2

用不确定性感知的简化模型,让离线强化学习更安全高效

Uncertainty-Aware Rank-One MIMO Q Network Framework for Accelerated Offline Reinforcement Learning

  • 通过量化数据不确定性并融入损失函数,指导模型学习更可靠的策略
  • 采用秩一多输入多输出结构,在单网络成本下实现类集成的不确定性感知能力
  • 在D4RL基准上达到顶尖性能,适合追求高效率离线强化学习的研究者

离线强化学习因其安全且易扩展的范式而备受关注,但其训练面临外分布(OOD)数据导致的外推误差问题。现有方法如惩罚OOD Q值或约束策略相似性,常因过度保守、对OOD数据刻画不准确或计算开销大而受限。本文提出一种不确定性感知的秩一多输入多输出(MIMO)Q网络框架,旨在充分挖掘OOD数据潜力的同时保证学习效率。该框架通过量化数据不确定性并将其融入训练损失,使策略最大化对应Q函数的置信下界。同时引入秩一MIMO结构,实现与网络集成相当的不确定性量化能力,但计算成本仅相当于单个网络。在D4RL基准上的大量实验表明,该框架在保持计算高效的同时达到当前最优性能,为缓解外推误差、提升离线强化学习效率提供了新路径。

原文摘要 · Abstract (English)

Offline reinforcement learning (RL) has garnered significant interest due to its safe and easily scalable paradigm. However, training under this paradigm presents its own challenge: the extrapolation error stemming from out-of-distribution (OOD) data. Existing methodologies have endeavored to address this issue through means like penalizing OOD Q-values or imposing similarity constraints on the learned policy and the behavior policy. Nonetheless, these approaches are often beset by limitations such as being overly conservative in utilizing OOD data, imprecise OOD data characterization, and significant computational overhead. To address these challenges, this paper introduces an Uncertainty-Aware Rank-One Multi-Input Multi-Output (MIMO) Q Network framework. The framework aims to enhance Offline Reinforcement Learning by fully leveraging the potential of OOD data while still ensuring efficiency in the learning process. Specifically, the framework quantifies data uncertainty and harnesses it in the training losses, aiming to train a policy that maximizes the lower confidence bound of the corresponding Q-function. Furthermore, a Rank-One MIMO architecture is introduced to model the uncertainty-aware Q-function, \TP{offering the same ability for uncertainty quantification as an ensemble of networks but with a cost nearly equivalent to that of a single network}. Consequently, this framework strikes a harmonious balance between precision, speed, and memory efficiency, culminating in improved overall performance. Extensive experimentation on the D4RL benchmark demonstrates that the framework attains state-of-the-art performance while remaining computationally efficient. By incorporating the concept of uncertainty quantification, our framework offers a promising avenue to alleviate extrapolation errors and enhance the efficiency of offline RL.

强化学习离线学习不确定性MIMO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。