arXiv:2603.15359cs.RO2026-03被引 7

让机器人提前预判行人动向并规划路径,提升复杂人群中的导航能力。

NavThinker: Action-Conditioned World Models for Coupled Prediction and Planning in Social Navigation

  • 用动作条件世界模型预测未来场景与行人轨迹
  • 在多机器人仿真中实现98.7%导航成功率,零样本迁移至真实环境
  • 适合需要安全交互的自主导航系统研发

社交导航要求机器人在动态人类环境中安全行动。有效行为需前瞻性思考:推理不同机器人动作下场景与行人的演化,而非仅依赖当前观测。这带来耦合预测-规划挑战,机器人动作与人类运动相互影响。为此,我们提出NavThinker,一个未来感知框架,将动作条件世界模型与在线策略强化学习结合。世界模型在Depth Anything V2块特征空间运行,进行自回归的未来场景几何与人类运动预测;多头解码器生成未来深度图与人类轨迹,产出与通行性及交互风险对齐的未来状态。关键在于,采用DD-PPO训练策略时,通过(i)将动作条件未来特征融合进当前观测嵌入,以及(ii)从预测行人轨迹获得社交奖励塑造,注入世界模型前瞻信号。在单/多机器人Social-HM3D上实验达领先水平导航成功率达98.7%,零样本迁移至Social-MP3D,并在Unitree Go2机器人上实现实体部署,验证泛化性与实用性。官网:https://hutslib.github.io/NavThinker。

原文摘要 · Abstract (English)

Social navigation requires robots to act safely in dynamic human environments. Effective behavior demands thinking ahead: reasoning about how the scene and pedestrians evolve under different robot actions rather than reacting to current observations alone. This creates a coupled prediction-planning challenge, where robot actions and human motion mutually influence each other. To address this challenge, we propose NavThinker, a future-aware framework that couples an action-conditioned world model with on-policy reinforcement learning. The world model operates in the Depth Anything V2 patch feature space and performs autoregressive prediction of future scene geometry and human motion; multi-head decoders then produce future depth maps and human trajectories, yielding a future-aware state aligned with traversability and interaction risk. Crucially, we train the policy with DD-PPO while injecting world-model think-ahead signals via: (i) action-conditioned future features fused into the current observation embedding and (ii) social reward shaping from predicted human trajectories. Experiments on single- and multi-robot Social-HM3D show state-of-the-art navigation success, with zero-shot transfer to Social-MP3D and real-world deployment on a Unitree Go2, validating generalization and practical applicability. Webpage: https://hutslib.github.io/NavThinker.

社交导航世界模型强化学习多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。