arXiv:2606.10495cs.RO2026-06

让视觉语言动作模型学会基于已有认知采取安全社交行为

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models

论文配图:Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models
图 1 · 摘自论文原文
  • 通过两阶段无标注微调,将内部社会特征与动作输出对齐
  • 在真实场景中近撞事件减少86.4%,反事实判断准确率达93%
  • 适合需安全社交导航的机器人部署,尤其关注行为可解释性

安全社交导航要求机器人区分行人与普通障碍物,并在危险发生前作出反应。我们发现预训练的视觉-语言-动作(VLA)模型已在内部表征中编码了行人-物体区分和未来碰撞信号,但行为克隆无法将这些信号转化为适当的社会行为。为此,我们提出SALSA——一种两阶段无标注后训练框架:(1) 社会行为对齐将中间层社会特征与动作头对齐,并在反事实人-物场景对上训练,以打破视觉显著性捷径;(2) 时间安全对齐提供自动生成的未来风险监督,实现前瞻式碰撞规避。在SCAND数据集和真实世界部署中,SALSA使近撞事件减少86.4%,社会反事实准确率从53%提升至93%,表明可通过更好对齐潜在表征与动作生成,让预训练VLA策略实现更安全的社交导航。

原文摘要 · Abstract (English)

Safe social navigation requires robots to distinguish people from ordinary obstacles and to react before danger becomes imminent. We show that pretrained Vision-Language-Action (VLA) models already encode pedestrian-object distinctions and future collision signals in their internal representations, but behavior cloning fails to translate these signals into socially appropriate actions. To address this mismatch, we propose SALSA, a two-stage annotation-free post-training framework: (1) social behavioral alignment bridges intermediate-layer social features to the action head and trains on counterfactual human-object scene pairs to break visual saliency shortcuts; (2) temporal safety alignment provides automatically generated future-risk supervision to enable anticipatory collision avoidance. On SCAND and real-world deployment, SALSA reduces near-collisions by 86.4% and improves social counterfactual accuracy from 53% to 93%, demonstrating that safer social navigation can be achieved by teaching VLA policies to act on representations they already possess. These results show that pretrained VLA policies can be adapted for safer social navigation by better aligning their latent representations with action generation.

机器人导航多模态对齐安全决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。