arXiv:2411.07760cs.LGcs.AI2024-11中稿 · ance被引 2

用空间量化提升复杂导航中的离线强化学习性能

Navigation with QPHIL: Quantizing Planner for Hierarchical Implicit Q-Learning

  • 基于变换器的分层规划,通过空间量化简化低层策略
  • 在长距离导航任务中达到当前最优效果
  • 适合研究离线强化学习与智能体导航的学者

离线强化学习已成为复杂导航任务中行为建模的有力替代方案,但其面临信号噪声比低的问题,即错误的价值估计导致策略更新偏差。现有方法表明,分层离线强化学习能将高层路径规划与低层路径追踪解耦。本文提出一种基于变换器的新型分层方法,利用空间的可学习量化机制,使低层策略仅依赖区域条件,简化了规划过程,将其转化为离散自回归预测。这种区域级推理支持显式轨迹拼接,而非依赖有噪价值函数的隐式拼接。结合近期离线强化学习进展,该方法在复杂长距离导航环境中取得当前最优表现。

原文摘要 · Abstract (English)

Offline Reinforcement Learning (RL) has emerged as a powerful alternative to imitation learning for behavior modeling in various domains, particularly in complex navigation tasks. An existing challenge with Offline RL is the signal-to-noise ratio, i.e. how to mitigate incorrect policy updates due to errors in value estimates. Towards this, multiple works have demonstrated the advantage of hierarchical offline RL methods, which decouples high-level path planning from low-level path following. In this work, we present a novel hierarchical transformer-based approach leveraging a learned quantizer of the space. This quantization enables the training of a simpler zone-conditioned low-level policy and simplifies planning, which is reduced to discrete autoregressive prediction. Among other benefits, zone-level reasoning in planning enables explicit trajectory stitching rather than implicit stitching based on noisy value function estimates. By combining this transformer-based planner with recent advancements in offline RL, our proposed approach achieves state-of-the-art results in complex long-distance navigation environments.

强化学习导航分层规划离线学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。