arXiv:2604.01659cs.RO2026-04被引 2

让机器人听懂人话,自动完成城市导航

AURA: Multimodal Shared Autonomy for Real-World Urban Navigation

  • 将导航拆分为人类指令与AI控制,降低操作负担
  • 实测减少44%以上人工接管次数,提升稳定性
  • 适合需要人机协作的自动驾驶、服务机器人场景

复杂城市环境中的长时程导航严重依赖持续的人类操作,导致疲劳、效率下降和安全隐患。共享自主通过视觉语言AI与人类操作员协同控制移动机器,提供有效解决方案。然而现有方法要求人机在同一动作空间操作,认知负担高。我们提出辅助城市机器人自主系统(AURA),一种多模态框架,将城市导航分解为高层人类指令与底层AI控制。AURA引入空间感知指令编码器,对齐人类指令与视觉空间上下文。为支持训练,构建了大规模数据集MM-CoS,包含远程操控与视觉语言描述。仿真与真实世界实验表明,AURA能有效遵循人类指令,降低人工操作负荷,提升导航稳定性,并支持在线适应。在相似接管条件下,共享自主框架使接管频率降低超44%。更多细节与演示视频见项目页面。

原文摘要 · Abstract (English)

Long-horizon navigation in complex urban environments relies heavily on continuous human operation, which leads to fatigue, reduced efficiency, and safety concerns. Shared autonomy, where a Vision-Language AI agent and a human operator collaborate on maneuvering the mobile machine, presents a promising solution to address these issues. However, existing shared autonomy methods often require humans and AI to operate within the same action space, leading to high cognitive overhead. We present Assistive Urban Robot Autonomy (AURA), a new multi-modal framework that decomposes urban navigation into high-level human instruction and low-level AI control. AURA incorporates a Spatial-Aware Instruction Encoder to align various human instructions with visual and spatial context. To facilitate training, we construct MM-CoS, a large-scale dataset comprising teleoperation and vision-language descriptions. Experiments in simulation and the real world demonstrate that AURA effectively follows human instructions, reduces manual operation effort, and improves navigation stability, while enabling online adaptation. Moreover, under similar takeover conditions, our shared autonomy framework reduces the frequency of takeovers by more than 44%. Demo video and more detail are provided in the project page.

人机协作城市导航多模态共享自主

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。