arXiv:2607.09336cs.LGcs.AI2026-07

用一键式快捷模型提升离线强化学习的规划效率

Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning

论文配图:Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning
图 1 · 摘自论文原文
  • 单阶段训练快捷轨迹模型,替代传统两阶段教学
  • 支持一步或多步推理,生成速度比扩散模型快10倍以上
  • 适合需要快速生成轨迹的机器人控制场景

基于扩散模型的轨迹规划在离线强化学习中表现优异,但其迭代去噪过程推理成本高。一致性规划虽减少采样步骤,却依赖耗时且不稳定的两阶段教师-学生蒸馏。我们提出快捷轨迹规划(STP),一种基于模型的离线强化学习框架,引入条件快捷轨迹模型作为高效轨迹生成器。STP采用单阶段训练,通过步长条件控制实现可调的一步或少步推理,并利用融合可行性修正的评判器筛选候选计划。在标准D4RL基准测试中,涵盖运动、导航、操作和灵巧控制任务,STP在保持强性能的同时,显著简化了训练流程,实现快速生成规划。

原文摘要 · Abstract (English)

Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling steps, yet they typically rely on a two-stage teacher--student distillation pipeline that increases training cost and may introduce instability. We propose Shortcut Trajectory Planning (STP), an offline model-based reinforcement learning framework that incorporates shortcut models as efficient trajectory generators. STP trains a conditional shortcut trajectory model in a single stage, supports adjustable one-step and few-step inference through step-size conditioning, and selects candidate plans using a critic augmented with feasibility-aware correction. Across standard D4RL benchmarks, including locomotion, navigation, manipulation, and dexterous control tasks, STP achieves strong performance while simplifying the training pipeline for fast generative planning.

轨迹规划强化学习高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。