arXiv:2507.11001cs.ROcs.CV2025-07被引 10

让机器人像专家一样自适应调参,提升复杂场景下的导航能力。

Learning to Tune Like an Expert: Interpretable and Scene-Aware Navigation via MLLM Reasoning and CVAE-Based Adaptation

  • 用大语言模型理解场景,结合变分自编码器自动调优导航参数。
  • 在多个规划器和真实场景中实现人类水平的导航性能。
  • 适合智能轮椅、服务机器人等需要社会感知的移动系统使用。

服务机器人在多样且动态的环境中部署日益增多,其物理布局与社交情境随时间和地点不断变化。传统依赖固定参数的导航系统难以跨场景泛化,导致性能下降与社会接受度降低。尽管近期方法采用强化学习改进传统规划器,但因泛化能力差及仿真多样性不足,难以实现有效的模拟到现实迁移。为此,我们提出LE-Nav框架,通过多模态大语言模型推理与条件变分自编码器,实现可解释且场景感知的导航参数自适应调节。为实现零样本场景理解,采用单样本示例与思维链提示策略;条件变分自编码器则建立自然语言指令与导航超参数间的映射关系,支持专家级调参。实验表明,LE-Nav可在多种规划器与场景中生成达到人类水平的超参数配置。真实世界导航测试与智能轮椅平台用户研究表明,其在成功率、效率、安全性和舒适性等量化指标上优于现有方法,同时获得更高的主观安全性与社会接受度评分。代码已开源:https://github.com/Cavendish518/LE-Nav。

原文摘要 · Abstract (English)

Service robots are increasingly deployed in diverse and dynamic environments, where both physical layouts and social contexts change over time and across locations. In these unstructured settings, conventional navigation systems that rely on fixed parameters often fail to generalize across scenarios, resulting in degraded performance and reduced social acceptance. Although recent approaches have leveraged reinforcement learning to enhance traditional planners, these methods often fail in real-world deployments due to poor generalization and limited simulation diversity, which hampers effective sim-to-real transfer. To tackle these issues, we present LE-Nav, an interpretable and scene-aware navigation framework that leverages multi-modal large language model reasoning and conditional variational autoencoders to adaptively tune planner hyperparameters. To achieve zero-shot scene understanding, we utilize one-shot exemplars and chain-of-thought prompting strategies. Additionally, a conditional variational autoencoder captures the mapping between natural language instructions and navigation hyperparameters, enabling expert-level tuning. Experiments show that LE-Nav can generate hyperparameters achieving human-level tuning across diverse planners and scenarios. Real-world navigation trials and a user study on a smart wheelchair platform demonstrate that it outperforms state-of-the-art methods on quantitative metrics such as success rate, efficiency, safety, and comfort, while receiving higher subjective scores for perceived safety and social acceptance. Code is available at https://github.com/Cavendish518/LE-Nav.

机器人导航大模型应用自适应调参

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。