提出量化扭曲机制,让离线强化学习更智能地保守。
Distorted Distributional Policy Evaluation for Offline Reinforcement Learning
- 用量化扭曲实现非均匀悲观,根据数据多少动态调整保守程度
- 在多个基准上优于传统均匀悲观方法,提升策略性能
- 适合追求离线强化学习稳定性和泛化能力的研究者
尽管分布式强化学习(DRL)在在线设置中表现优异,但在离线场景中仍受限。我们假设现有离线DRL方法因对回报分位数进行均匀低估而效果不佳,这种均匀悲观会导致估值过于保守,进而影响泛化能力。为此,本文提出量化扭曲新概念,通过根据支持数据的可用性调整悲观程度,实现非均匀悲观。该方法基于理论分析,并经实证验证,在多个基准测试中表现优于均匀悲观策略,显著提升离线强化学习性能。
原文摘要 · Abstract (English)
While Distributional Reinforcement Learning (DRL) methods have demonstrated strong performance in online settings, its success in offline scenarios remains limited. We hypothesize that a key limitation of existing offline DRL methods lies in their approach to uniformly underestimate return quantiles. This uniform pessimism can lead to overly conservative value estimates, ultimately hindering generalization and performance. To address this, we introduce a novel concept called quantile distortion, which enables non-uniform pessimism by adjusting the degree of conservatism based on the availability of supporting data. Our approach is grounded in theoretical analysis and empirically validated, demonstrating improved performance over uniform pessimism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。