arXiv:2607.10991cs.ROcs.AI2026-07

让机器人在关键社交场景下临时调用大模型思考,兼顾实时性与理解力。

Think When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies

论文配图:Think When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies
图 1 · 摘自论文原文
  • RL策略主控日常导航,有人进入敏感区时触发预训练视觉语言模型推理
  • 在Social-MP3D上任务成功率提升20%,个人空间侵犯减少显著
  • 适合需要高安全性和自然交互的现实场景机器人部署

随着移动机器人日益融入日常人类环境,社交导航成为保障人类舒适度、安全与信任的关键。尽管强化学习(RL)导航策略具备实时推理与反应能力,但缺乏灵活的语义推理能力,难以泛化至复杂社交场景。近期研究开始用视觉语言模型(VLM)替代RL策略以增强语义与社交理解,但其高计算开销和慢推理速度仍是实时部署的主要障碍。为此,我们提出HUMA(Hybrid Understanding for Multi-modal social Navigation),一种动态平衡RL策略计算效率与VLM深度语义理解的混合架构。该方法使用反应式RL策略处理低密度、常规导航任务,当人进入机器人敏感区域(如靠近范围)时,条件性调用预训练的高层VLM进行推理。我们在Social-MP3D与Social-HM3D基准上评估HUMA,分别实现20%和3%的任务成功率提升,同时显著降低个人空间侵犯与人类碰撞。消融实验验证了各组件有效性,真实机器人(Mirokaï)部署进一步证明了该方法的实用性。

原文摘要 · Abstract (English)

As mobile robots become more integrated into everyday human environments, social robot navigation is becoming essential for ensuring human comfort, safety, and trust. While reinforcement learning (RL) navigation policies provide the fast inference and reactive behavior necessary for real-time deployment, they still lack flexible semantic reasoning capabilities and often fail to generalize to complex social scenarios. Recent approaches have increasingly turned to vision-language models (VLMs) in place of RL policies to improve semantic and social reasoning in robot navigation. Nevertheless, their high computational cost and slow inference remain major barriers to real-time deployment. To overcome these limitations, we introduce HUMA (Hybrid Understanding for Multi-modal social Navigation), a hybrid architecture that dynamically balances the computational efficiency of RL policies with the deep semantic understanding of VLMs. Our approach uses a reactive RL policy to handle low-density, routine navigation tasks, while conditioning it on a post-trained high-level VLM when a human enters sensitive situations, such as the robot's proximity zone. We evaluate HUMA on the Social-MP3D and Social-HM3D benchmarks, where it achieves task success improvements of 20% and 3%, respectively, while significantly reducing personal space violations and human collisions against state-of-the-art baselines. Extensive ablation studies validate each architectural component, and real-world deployment on the Mirokaï mobile robot further demonstrates the practical viability of our approach.

社交导航强化学习视觉语言模型混合架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。