arXiv:2605.08202cs.LGcs.AI2026-05中稿 · ICLR被引 1

用扩散模型精准识别离群动作,提升离线强化学习安全性与探索效率

Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning

论文配图:Beyond Penalization: Diffusion-based Out-of-Distribution Detection and Selective Regularization in Offline Reinforcement Learning
图 1 · 摘自论文原文
  • 基于双扩散模型构建行为策略与状态分布,用去噪误差判别离群动作
  • 在优化中区分有益与有害离群动作,使性能超越现有方法,尤其在低质量数据上
  • 理论证明收敛性与近似最优性,适合数据质量差的离线强化学习场景

离线强化学习面临对离群动作价值高估的核心挑战。现有方法通过统一惩罚未见样本缓解此问题,但难以准确识别离群动作,且可能抑制有益探索。尽管已有方法尝试利用不同属性区分离群样本,但通常依赖严苛的数据分布假设,判别能力有限。为此,本文提出DOSER(Diffusion-based OOD Detection and Selective Regularization),一种超越均匀惩罚的新框架。DOSER训练两个扩散模型以捕捉行为策略和状态分布,采用单步去噪重建误差作为可靠的离群检测指标。在策略优化过程中,通过预测转移结果进一步区分有益与有害离群动作,选择性抑制风险动作,同时鼓励高潜力探索。理论上,我们证明DOSER是$γ$-收缩的,因此具有唯一不动点且价值估计有界。此外,我们提供了相对于最优策略的渐近性能保证,考虑模型近似与离群检测误差。在广泛离线强化学习基准测试中,DOSER持续优于先前方法,尤其在次优数据集上表现显著。

原文摘要 · Abstract (English)

Offline reinforcement learning (RL) faces a critical challenge of overestimating the value of out-of-distribution (OOD) actions. Existing methods mitigate this issue by penalizing unseen samples, yet they fail to accurately identify OOD actions and may suppress beneficial exploration beyond the behavioral support. Although several methods have been proposed to differentiate OOD samples with distinct properties, they typically rely on restrictive assumptions about the data distribution and remain limited in discrimination ability. To address this problem, we propose DOSER (Diffusion-based OOD Detection and Selective Regularization), a novel framework that goes beyond uniform penalization. DOSER trains two diffusion models to capture the behavior policy and state distribution, using single-step denoising reconstruction error as a reliable OOD indicator. During policy optimization, it further distinguishes between beneficial and detrimental OOD actions by evaluating predicted transitions, selectively suppressing risky actions while encouraging exploration of high-potential ones. Theoretically, we prove that DOSER is a $γ$-contraction and therefore admits a unique fixed point with bounded value estimates. We further provide an asymptotic performance guarantee relative to the optimal policy under model approximation and OOD detection errors. Across extensive offline RL benchmarks, DOSER consistently attains superior performance to prior methods, especially on suboptimal datasets.

强化学习离线学习扩散模型异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。