让机器人像人一样思考,智能理解用户意图并导航找物
CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking
- 引入双思维系统,快速决策+持续学习优化
- 在ProcThor数据集上导航准确率提升15%
- 适合需要理解隐含指令的智能交互场景
移动机器人需在未知非结构化环境中满足人类需求。需求驱动导航(DDN)使机器人能基于隐含的人类意图识别并定位物体,即使物体位置未知。然而,传统数据驱动的DDN方法依赖预收集数据训练模型,泛化能力受限。本文提出基于视觉语言模型(VLM)的CogDDN框架,模拟人类认知与学习机制,融合快速与慢速思维系统,选择性识别关键目标以满足用户需求。通过语义对齐检测到的物体与指令,确定目标对象。同时引入双过程决策模块:启发式过程实现快速高效决策,分析过程则分析过往错误、积累至知识库并持续改进性能。思维链(CoT)推理强化决策过程。在AI2Thor仿真器和ProcThor数据集上的闭环评估显示,相比单视角相机方法,CogDDN导航准确率提升15%,显著提升导航精度与适应性。
原文摘要 · Abstract (English)
Mobile robots are increasingly required to navigate and interact within unknown and unstructured environments to meet human demands. Demand-driven navigation (DDN) enables robots to identify and locate objects based on implicit human intent, even when object locations are unknown. However, traditional data-driven DDN methods rely on pre-collected data for model training and decision-making, limiting their generalization capability in unseen scenarios. In this paper, we propose CogDDN, a VLM-based framework that emulates the human cognitive and learning mechanisms by integrating fast and slow thinking systems and selectively identifying key objects essential to fulfilling user demands. CogDDN identifies appropriate target objects by semantically aligning detected objects with the given instructions. Furthermore, it incorporates a dual-process decision-making module, comprising a Heuristic Process for rapid, efficient decisions and an Analytic Process that analyzes past errors, accumulates them in a knowledge base, and continuously improves performance. Chain of Thought (CoT) reasoning strengthens the decision-making process. Extensive closed-loop evaluations on the AI2Thor simulator with the ProcThor dataset show that CogDDN outperforms single-view camera-only methods by 15\%, demonstrating significant improvements in navigation accuracy and adaptability. The project page is available at https://yuehaohuang.github.io/CogDDN/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。