将视觉语言模型的社会理解能力提炼为导航注意力图,提升机器人社交导航成功率。
ViLAM: Distilling Vision-Language Reasoning into Attention Maps for Social Robot Navigation
- 在注意力层进行知识蒸馏,对齐大模型与动作模型的注意力图。
- 实测在真实机器人上成功率达14.2%~50%提升,显著优于现有方法。
- 适合关注机器人社交导航与多模态感知融合的研究者。
我们提出ViLAM,一种将大型视觉语言模型(VLM)中的视觉-语言推理能力,蒸馏为用于社会合规机器人导航的空间注意力图的新方法。不同于依赖专家示范或人工标注数据集的传统方法,ViLAM在中间层表示(注意力)层面进行知识蒸馏与微调,通过将预训练视觉-动作模型的注意力图与由大型VLM生成的社会引导注意力图对齐。这些蒸馏后的注意力图突出场景中的关键导航区域,并作为运动规划中的社会感知空间代价地图。为此,我们设计了一种新型注意力级蒸馏损失,融合双源知识,生成具有增强社会意识的增强注意力图。这些优化后的注意力图被用于社会感知局部规划器中作为可通行性代价图。我们在Husky轮式机器人上进行真实世界实验,验证了该方法的有效性,结果显示相较于现有方法,成功率提升了14.2%至50%。
原文摘要 · Abstract (English)
We introduce ViLAM, a novel method for distilling vision-language reasoning from large Vision-Language Models (VLMs) into spatial attention maps for socially compliant robot navigation. Unlike traditional methods that rely on expert demonstrations or human-annotated datasets, ViLAM performs knowledge distillation and fine-tuning at the intermediate layer representation (attention) level by aligning attention maps from a pretrained vision-action model with socially guided attention maps derived from a large VLM. These distilled attention maps highlight key navigational regions in a scene and serve as socially informed spatial cost maps for motion planning. To achieve this, we introduce a novel attention-level distillation loss that fuses knowledge from both sources, generating augmented attention maps with enhanced social awareness. These refined attention maps are then used as a traversability costmap within a socially aware local planner for navigation. We validate our approach through real-world experiments on a Husky wheeled robot, and demonstrate 14.2% - 50% improvements in success rate over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。