用视觉语言模型自动调参,让机器人在复杂环境更安全地导航
APPLV: Adaptive Planner Parameter Learning from Vision-Language-Action Model
- 用预训练模型预测经典规划器参数,不直接输出动作
- 在仿真和真实机器人上均提升导航性能与泛化能力
- 适合需要高安全性和强适应性的移动机器人场景
在高度受限环境中实现自主导航对移动机器人仍是挑战。传统导航方法虽具安全性保障,但需针对环境调参;端到端学习省去调参却难以精确控制。近期方法尝试自动化调参并保留传统系统的安全性,但在未见环境中的泛化能力仍不足。视觉-语言-动作(VLA)模型凭借基础模型的场景理解能力展现出潜力,但仍面临导航任务中精确控制与推理延迟的问题。本文提出从视觉-语言-动作模型中自适应学习规划器参数的方法(APPLV)。不同于传统VLA模型直接输出动作,APPLV利用预训练视觉-语言模型加回归头,预测配置经典规划器的参数。我们设计两种训练策略:基于采集导航轨迹的监督微调,以及进一步优化导航性能的强化学习微调。在模拟基准自主机器人导航(BARN)数据集及物理机器人实验中评估APPLV,结果表明其在导航性能和未见环境泛化方面均优于现有方法。
原文摘要 · Abstract (English)
Autonomous navigation in highly constrained environments remains challenging for mobile robots. Classical navigation approaches offer safety assurances but require environment-specific parameter tuning; end-to-end learning bypasses parameter tuning but struggles with precise control in constrained spaces. To this end, recent robot learning approaches automate parameter tuning while retaining classical systems' safety, yet still face challenges in generalizing to unseen environments. Recently, Vision-Language-Action (VLA) models have shown promise by leveraging foundation models' scene understanding capabilities, but still struggle with precise control and inference latency in navigation tasks. In this paper, we propose Adaptive Planner Parameter Learning from Vision-Language-Action Model (\textsc{applv}). Unlike traditional VLA models that directly output actions, \textsc{applv} leverages pre-trained vision-language models with a regression head to predict planner parameters that configure classical planners. We develop two training strategies: supervised learning fine-tuning from collected navigation trajectories and reinforcement learning fine-tuning to further optimize navigation performance. We evaluate \textsc{applv} across multiple motion planners on the simulated Benchmark Autonomous Robot Navigation (BARN) dataset and in physical robot experiments. Results demonstrate that \textsc{applv} outperforms existing methods in both navigation performance and generalization to unseen environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。