arXiv:2506.00895cs.LGcs.AI2025-06NeurIPS被引 15

通过拼接轨迹提升扩散规划器的长程推理能力

State-Covering Trajectory Stitching for Diffusion Planners

  • 用时序保距的隐空间表示环境,逐步拼接轨迹段
  • 在多样性和新颖性引导下扩展隐空间覆盖范围
  • 显著提升离线强化学习中长程任务的泛化性能

基于扩散的生成模型正成为强化学习中长程规划的强大工具,尤其适用于离线数据集。然而,其性能受限于训练数据的质量与多样性,难以泛化到训练分布之外的任务或更长的规划时序。为此,我们提出状态覆盖轨迹拼接(SCoTS),一种无奖励的轨迹增强方法,通过增量拼接短轨迹段,系统性地生成多样化且更长的轨迹。SCoTS首先学习一个保持时序结构的隐空间表示,随后在方向探索和新颖性引导下迭代拼接轨迹段,有效覆盖并扩展该隐空间。实验表明,SCoTS显著提升了扩散规划器在需拼接与长程推理的离线目标条件基准上的表现,且生成的增强轨迹也大幅改善了多种环境中主流离线目标条件强化学习算法的性能。

原文摘要 · Abstract (English)

Diffusion-based generative models are emerging as powerful tools for long-horizon planning in reinforcement learning (RL), particularly with offline datasets. However, their performance is fundamentally limited by the quality and diversity of training data. This often restricts their generalization to tasks outside their training distribution or longer planning horizons. To overcome this challenge, we propose State-Covering Trajectory Stitching (SCoTS), a novel reward-free trajectory augmentation method that incrementally stitches together short trajectory segments, systematically generating diverse and extended trajectories. SCoTS first learns a temporal distance-preserving latent representation that captures the underlying temporal structure of the environment, then iteratively stitches trajectory segments guided by directional exploration and novelty to effectively cover and expand this latent space. We demonstrate that SCoTS significantly improves the performance and generalization capabilities of diffusion planners on offline goal-conditioned benchmarks requiring stitching and long-horizon reasoning. Furthermore, augmented trajectories generated by SCoTS significantly improve the performance of widely used offline goal-conditioned RL algorithms across diverse environments.

扩散模型长程规划轨迹增强离线RL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。