arXiv:2512.00049cs.ROcs.AI2025-12综述被引 9

用深度强化学习让机器人在人多环境里自然行走,既安全又不扰人。

Socially aware navigation for mobile robots: a survey on deep reinforcement learning approaches

  • 结合深度强化学习与社交规范建模,让机器人学会人类行为习惯。
  • 现有方法显著提升安全性与接受度,但评估标准不统一制约发展。
  • 适合研究人机交互、智能导航的学者和工程师参考。

社会感知导航是机器人学中快速发展的领域,使机器人能够在人类环境中移动时遵循隐含的社会规范。深度强化学习(DRL)的兴起加速了导航策略的发展,使机器人能在实现目标的同时融入社会惯例。本文全面综述了基于DRL的社会感知导航方法,重点涵盖距离规范(proxemics)、人类舒适度、行为自然性、轨迹与意图预测等关键方面,以提升机器人在人类环境中的互动质量。文章深入分析了基于价值、策略及演员-评论家框架的DRL算法,以及前馈、循环、卷积、图神经网络与变换器等神经网络架构在增强代理学习与表征能力中的应用。同时,探讨了评价机制,包括评估指标、基准数据集、仿真环境,以及模拟到现实迁移的持续挑战。对比分析表明,尽管DRL显著提升了安全性和人类接受度,但该领域仍面临评估标准不统一、缺乏标准化社交度量、计算负担重导致可扩展性差,以及仿真到真实硬件部署困难等问题。未来进展依赖于融合多种方法的混合范式,以及兼顾技术效率与以人为中心评估的基准体系。

原文摘要 · Abstract (English)

Socially aware navigation is a fast-evolving research area in robotics that enables robots to move within human environments while adhering to the implicit human social norms. The advent of Deep Reinforcement Learning (DRL) has accelerated the development of navigation policies that enable robots to incorporate these social conventions while effectively reaching their objectives. This survey offers a comprehensive overview of DRL-based approaches to socially aware navigation, highlighting key aspects such as proxemics, human comfort, naturalness, trajectory and intention prediction, which enhance robot interaction in human environments. This work critically analyzes the integration of value-based, policy-based, and actor-critic reinforcement learning algorithms alongside neural network architectures, such as feedforward, recurrent, convolutional, graph, and transformer networks, for enhancing agent learning and representation in socially aware navigation. Furthermore, we examine crucial evaluation mechanisms, including metrics, benchmark datasets, simulation environments, and the persistent challenges of sim-to-real transfer. Our comparative analysis of the literature reveals that while DRL significantly improves safety, and human acceptance over traditional approaches, the field still faces setback due to non-uniform evaluation mechanisms, absence of standardized social metrics, computational burdens that limit scalability, and difficulty in transferring simulation to real robotic hardware applications. We assert that future progress will depend on hybrid approaches that leverage the strengths of multiple approaches and producing benchmarks that balance technical efficiency with human-centered evaluation.

机器人导航强化学习人机交互社会规范

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。