arXiv:2608.12917cs.LGcs.RO2026-08

用人体距离理论优化机器人社交导航,让机器人更懂人际空间

Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling

论文配图:Towards Socially Compliant Navigation in Deep Reinforcement Learning via Proxemics-Based Reward Modeling
图 1 · 摘自论文原文
  • 基于霍尔的人际距离理论构建密度场奖励信号
  • 在多种人群密度下提升社交合规性指标,导航效率不降
  • 适合需要自然互动的机器人场景,如商场导览

在拥挤环境中的有效机器人导航对实际应用至关重要。尽管近期深度强化学习(DRL)方法提升了复杂环境下的导航表现,但多聚焦任务目标,忽视社交合规性。本文提出一种基于邻近关系(proxemics)的奖励建模方法,为DRL社交导航提供密集且可解释的社会学习信号,同时保持导航效率。该方法将每个人类的个人空间建模为基于霍尔邻近关系理论的径向高斯混合场,并在机器人视域范围内计算以机器人为中心的局部代价。我们将该奖励机制整合至主流DRL导航方法,在多种人群场景、奖励基线及人群密度下进行仿真评估,使用导航与社会性指标双重验证。结果表明,该奖励在保持与对比模型相当导航性能的同时,持续显著改善社会性指标。

原文摘要 · Abstract (English)

Developing effective robot navigation methods in crowded environments is essential for real-world applications. Although recent deep reinforcement learning (DRL) methods have improved navigation performance in crowded environments, they often focus primarily on task-centric objectives and underrepresent social compliance objectives. In this paper, we introduce a novel proxemics-based reward formulation for DRL social navigation that provides a dense, interpretable social learning signal while maintaining navigation efficiency. Our approach models each human's personal space as a radial Gaussian-mixture field derived from Hall's proxemics theory and computes a robot-centric local cost over the robot's field of view. We integrate the proposed reward into established DRL navigation methods and evaluate it in simulation across multiple crowd scenarios, reward baselines, and crowd densities using both navigation metrics and social metrics. Results show that the proposed reward consistently improves social metrics in simulation while maintaining competitive navigation performance relative to the compared reward models.

机器人导航强化学习社交合规邻近关系

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。