arXiv:2510.00466cs.ROcs.AI2025-10

用因果Transformer融合轨迹预测,让机器人导航更稳更安全。

Integrating Offline Pre-Training with Online Fine-Tuning: A Reinforcement Learning Approach for Robot Social Navigation

  • 用时空融合模型实时预测到达目标的收益(RTG)
  • 实验显示成功率更高、碰撞率更低,优于现有方法
  • 适合需要实时适应人群的机器人导航场景

离线强化学习在机器人社交导航中展现出潜力,但行人行为的不确定性及训练阶段环境交互有限,常导致探索不足和离线与在线分布差异。本文提出一种新型离线到在线微调强化学习算法,将返程收益(Return-to-Go, RTG)预测融入因果Transformer架构。通过联合建模时间序列行人运动模式与空间人群动态,设计时空融合模型实现实时精准的RTG估计,缓解分布偏移问题。同时构建混合离线-在线经验采样机制,稳定微调过程中的策略更新,实现预训练知识与实时适应的平衡。大量仿真环境实验表明,该方法在成功率和碰撞率上均优于当前最优基线,验证了其在提升导航策略鲁棒性与适应性方面的有效性。本工作为真实场景下更可靠、自适应的机器人导航系统提供新路径。

原文摘要 · Abstract (English)

Offline reinforcement learning (RL) has emerged as a promising framework for addressing robot social navigation challenges. However, inherent uncertainties in pedestrian behavior and limited environmental interaction during training often lead to suboptimal exploration and distributional shifts between offline training and online deployment. To overcome these limitations, this paper proposes a novel offline-to-online fine-tuning RL algorithm for robot social navigation by integrating Return-to-Go (RTG) prediction into a causal Transformer architecture. Our algorithm features a spatiotem-poral fusion model designed to precisely estimate RTG values in real-time by jointly encoding temporal pedestrian motion patterns and spatial crowd dynamics. This RTG prediction framework mitigates distribution shift by aligning offline policy training with online environmental interactions. Furthermore, a hybrid offline-online experience sampling mechanism is built to stabilize policy updates during fine-tuning, ensuring balanced integration of pre-trained knowledge and real-time adaptation. Extensive experiments in simulated social navigation environments demonstrate that our method achieves a higher success rate and lower collision rate compared to state-of-the-art baselines. These results underscore the efficacy of our algorithm in enhancing navigation policy robustness and adaptability. This work paves the way for more reliable and adaptive robotic navigation systems in real-world applications.

机器人导航强化学习因果模型实时适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。