arXiv:2509.21942cs.LG2025-09NeurIPS被引 2

基于结构信息自适应构建扩散层次,提升长程离线强化学习性能

Structural Information-based Hierarchical Diffusion for Offline Reinforcement Learning

  • 根据轨迹结构信息动态构建多尺度扩散层次
  • 在复杂任务中超越现有方法,提升决策稳定性和泛化能力
  • 适合长时序、稀疏奖励的离线强化学习场景

基于生成式方法的扩散模型在建模离线强化学习数据集轨迹方面展现出巨大潜力,而分层扩散已被用于缓解长时程规划中的方差累积与计算挑战。然而,现有方法通常采用固定双层结构和单一预设时间尺度,限制了对多样化下游任务的适应性并降低决策灵活性。本文提出SIHD——一种基于结构信息的分层扩散框架,用于在长时程、稀疏奖励环境中实现高效稳定的离线策略学习。我们分析离线轨迹中嵌入的结构信息,自适应构建扩散层次,实现多时间尺度的灵活轨迹建模。不同于依赖局部子轨迹奖励预测的方法,我们通过量化每个状态社区的结构信息增益,将其作为对应扩散层的条件信号。为减少对离线数据集的过度依赖,引入结构熵正则项,鼓励探索低频状态,同时避免分布偏移带来的外推误差。在多个具有挑战性的离线强化学习任务上的实验表明,SIHD显著优于当前最优基线,在决策表现和跨场景泛化能力上均表现出色。

原文摘要 · Abstract (English)

Diffusion-based generative methods have shown promising potential for modeling trajectories from offline reinforcement learning (RL) datasets, and hierarchical diffusion has been introduced to mitigate variance accumulation and computational challenges in long-horizon planning tasks. However, existing approaches typically assume a fixed two-layer diffusion hierarchy with a single predefined temporal scale, which limits adaptability to diverse downstream tasks and reduces flexibility in decision making. In this work, we propose SIHD, a novel Structural Information-based Hierarchical Diffusion framework for effective and stable offline policy learning in long-horizon environments with sparse rewards. Specifically, we analyze structural information embedded in offline trajectories to construct the diffusion hierarchy adaptively, enabling flexible trajectory modeling across multiple temporal scales. Rather than relying on reward predictions from localized sub-trajectories, we quantify the structural information gain of each state community and use it as a conditioning signal within the corresponding diffusion layer. To reduce overreliance on offline datasets, we introduce a structural entropy regularizer that encourages exploration of underrepresented states while avoiding extrapolation errors from distributional shifts. Extensive evaluations on challenging offline RL tasks show that SIHD significantly outperforms state-of-the-art baselines in decision-making performance and demonstrates superior generalization across diverse scenarios.

离线RL扩散模型分层建模结构信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。