arXiv:2605.13321cs.RO2026-05

让机器人像人一样理解他人意图,实现安全社交导航。

HCSG: Human-Centric Semantic-Geometric Reasoning for Vision-Language Navigation

论文配图:HCSG: Human-Centric Semantic-Geometric Reasoning for Vision-Language Navigation
图 1 · 摘自论文原文
  • 融合语义与几何信息,实时预测人类动作和轨迹。
  • 在HA-VLNCE上成功率达74.3%,碰撞率降低34%。
  • 适合研究人机交互、智能导航的开发者参考。

视觉语言导航(VLN)在扩大数据与模型规模后取得了显著进展。然而,静态环境假设在真实室内场景中失效,机器人不可避免会遇到动态行人。现有方法通常仅将人视为基于隐式视觉线索的移动障碍物,缺乏对人类意图的显式推理或社会规范的遵守能力。为此,我们提出首个以人类为中心的框架HCSG,推动从被动避障转向主动理解人类行为。HCSG引入统一的人类理解模块,结合两项核心能力:(i) 几何预测,通过估计人体姿态与轨迹来预判未来运动动态;(ii) 语义解析,利用视觉语言模型(VLM)生成人类动作与意图的自然语言描述。这些语义-几何表示被融合至代理的拓扑地图中,用于指令条件下的路径规划。此外,引入社交距离损失,强制保持符合社会规范的交互距离。在HA-VLNCE基准上的大量实验表明,HCSG显著优于当前最优方法,成功率达74.3%,碰撞率下降34%。

原文摘要 · Abstract (English)

VLN has achieved remarkable progress by scaling data and model capacity. However, the assumption of a static environment breaks down in real-world indoor scenarios, where robots inevitably encounter dynamic pedestrians. Existing human-aware approaches typically treat humans merely as moving obstacles based on implicit visual cues, lacking the explicit reasoning required to interpret human intentions or maintain social norms. To address this, we propose HCSG, the first human-centric framework for VLN. This framework provides a robust foundation for safe, socially intelligent navigation in dynamic human-robot environments that shifts the paradigm from passive collision avoidance to active human behavior understanding. Specifically, HCSG introduces a unified Human Understanding Module that synergizes two key capabilities: (i) geometric forecasting, which predicts human pose and trajectory to anticipate future motion dynamics; and (ii) semantic interpretation, which leverages a Vision-Language Model (VLM) to generate natural language descriptions of human actions and intentions. These semantic-geometric representations are fused into the agent's topological map for instruction-conditioned planning. Furthermore, a social distance loss is introduced to enforce socially compliant interaction distances. Extensive experiments on the HA-VLNCE benchmark demonstrate that HCSG significantly outperforms state-of-the-art methods, achieving a 14% improvement in Success Rate and a 34% reduction in Collision Rate. Our project can be seen at https://haoxuanxu1024.github.io/HCSG/.

视觉导航人机交互语义理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。