arXiv:2602.02900cs.LGcs.AI2026-02被引 6

用隐空间能量模型约束策略,提升离线强化学习在分布外时的稳定性。

Manifold-Constrained Energy-Based Transition Models for Offline Reinforcement Learning

  • 通过隐空间扰动与朗之万动力学生成近曼德拉硬负样本,增强对分布外状态的敏感度。
  • 在标准控制任务上,相比基线提升多步动态保真度和归一化回报,尤其在稀疏数据下表现更好。
  • 适合研究离线强化学习中如何避免模型误差累积与过估计问题的学者。

基于模型的离线强化学习在分布偏移下易失效:策略优化会将模拟轨迹推向数据集支持较弱的状态-动作区域,导致模型误差累积并引发严重价值过估计。本文提出曼德拉约束的能量型转移模型(MC-ETM),利用流形投影-扩散负采样训练条件能量模型。MC-ETM学习下一状态的隐空间流形,通过扰动隐码并在隐空间中运行朗之万动力学生成近流形硬负样本,从而在数据集支持区域附近锐化能量景观,提升对微小分布外偏差的敏感性。在策略优化中,学习到的能量提供单一可靠性信号:当采样下一状态的最小能量超过阈值时截断轨迹;贝尔曼更新通过基于能量引导样本间Q值离散度的悲观惩罚实现稳定。我们通过混合悲观马尔可夫决策过程形式化了MC-ETM,并推导出保守性能界,分离了支持内评估误差与截断风险。实验表明,MC-ETM提升了多步动态保真度,在标准离线控制基准上取得更高归一化回报,尤其在非规则动态与稀疏数据覆盖条件下优势显著。

原文摘要 · Abstract (English)

Model-based offline reinforcement learning is brittle under distribution shift: policy improvement drives rollouts into state--action regions weakly supported by the dataset, where compounding model error yields severe value overestimation. We propose Manifold-Constrained Energy-based Transition Models (MC-ETM), which train conditional energy-based transition models using a manifold projection--diffusion negative sampler. MC-ETM learns a latent manifold of next states and generates near-manifold hard negatives by perturbing latent codes and running Langevin dynamics in latent space with the learned conditional energy, sharpening the energy landscape around the dataset support and improving sensitivity to subtle out-of-distribution deviations. For policy optimization, the learned energy provides a single reliability signal: rollouts are truncated when the minimum energy over sampled next states exceeds a threshold, and Bellman backups are stabilized via pessimistic penalties based on Q-value-level dispersion across energy-guided samples. We formalize MC-ETM through a hybrid pessimistic MDP formulation and derive a conservative performance bound separating in-support evaluation error from truncation risk. Empirically, MC-ETM improves multi-step dynamics fidelity and yields higher normalized returns on standard offline control benchmarks, particularly under irregular dynamics and sparse data coverage.

强化学习离线学习能量模型分布外检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。