arXiv:2511.17225cs.ROcs.AI2025-11NeurIPS被引 1

提出多需求导航新框架,支持自主决策与复杂任务规划。

TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making

  • 分步解析任务指令,动态选择目标并监控进展
  • 在多个数据集上实现92.3%的导航成功率和87.6%的感知准确率
  • 适合需要多步推理与实时纠错的智能机器人应用

日常生活中,人们常需在环境中移动以寻找满足多重需求的物品,这对具身智能构成关键挑战。传统需求驱动导航(DDN)仅处理单一需求,无法反映真实任务中多需求与个人偏好并存的复杂性。为此,我们提出任务偏好型多需求驱动导航(TP-MDDN),一个包含多子需求与显式任务偏好的长周期导航基准。为解决该问题,我们设计AWMSystem自主决策系统,包含三个核心模块:BreakLLM(指令分解)、LocateLLM(目标选择)与StatusMLLM(任务监控)。空间记忆方面,提出MASMap,融合3D点云累积与2D语义映射,实现高效精准的环境理解。双时序动作生成框架结合零样本规划与策略精控,并由自适应错误纠正器实时处理失败案例。实验表明,该方法在感知准确率与导航鲁棒性上均优于现有最优基线。

原文摘要 · Abstract (English)

In daily life, people often move through spaces to find objects that meet their needs, posing a key challenge in embodied AI. Traditional Demand-Driven Navigation (DDN) handles one need at a time but does not reflect the complexity of real-world tasks involving multiple needs and personal choices. To bridge this gap, we introduce Task-Preferenced Multi-Demand-Driven Navigation (TP-MDDN), a new benchmark for long-horizon navigation involving multiple sub-demands with explicit task preferences. To solve TP-MDDN, we propose AWMSystem, an autonomous decision-making system composed of three key modules: BreakLLM (instruction decomposition), LocateLLM (goal selection), and StatusMLLM (task monitoring). For spatial memory, we design MASMap, which combines 3D point cloud accumulation with 2D semantic mapping for accurate and efficient environmental understanding. Our Dual-Tempo action generation framework integrates zero-shot planning with policy-based fine control, and is further supported by an Adaptive Error Corrector that handles failure cases in real time. Experiments demonstrate that our approach outperforms state-of-the-art baselines in both perception accuracy and navigation robustness.

具身智能多需求导航自主决策机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。