用单步生成实现高速高质轨迹规划,显著提升离线强化学习效率。
Consistency Trajectory Planning: High-Quality and Efficient Trajectory Optimization for Offline Model-Based Reinforcement Learning
- 基于一致性轨迹模型,单步生成高质量轨迹,避免多步迭代。
- 在长时序目标导向任务中,性能优于现有扩散模型方法,推理速度提升120倍以上。
- 适合需要低延迟、高精度轨迹生成的工业级离线强化学习应用。
本文提出一致性轨迹规划(CTP),一种基于近期提出的一致性轨迹模型(CTM)的新型离线模型化强化学习方法,实现高效轨迹优化。尽管先前将扩散模型用于规划的方法表现出色,但常因迭代采样带来高昂计算开销。CTP 支持无需显著性能下降的快速单步轨迹生成。我们在 D4RL 基准上评估了 CTP,结果表明其在长时域、目标条件任务中持续优于现有扩散模型规划方法。特别地,CTP 在使用更少去噪步骤的情况下获得更高归一化回报,在推理时间上实现超过120倍的加速,充分展示了其在高性能、低延迟离线规划中的实用性与有效性。
原文摘要 · Abstract (English)
This paper introduces Consistency Trajectory Planning (CTP), a novel offline model-based reinforcement learning method that leverages the recently proposed Consistency Trajectory Model (CTM) for efficient trajectory optimization. While prior work applying diffusion models to planning has demonstrated strong performance, it often suffers from high computational costs due to iterative sampling procedures. CTP supports fast, single-step trajectory generation without significant degradation in policy quality. We evaluate CTP on the D4RL benchmark and show that it consistently outperforms existing diffusion-based planning methods in long-horizon, goal-conditioned tasks. Notably, CTP achieves higher normalized returns while using significantly fewer denoising steps. In particular, CTP achieves comparable performance with over $120\times$ speedup in inference time, demonstrating its practicality and effectiveness for high-performance, low-latency offline planning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。