arXiv:2602.05273cs.RO2026-02

让机器人理解模糊指令并自主找物执行,成功率超80%。

Affordance-Aware Interactive Decision-Making and Execution for Ambiguous Instructions

  • 双流架构:分决策与执行两路,结合视觉语言推理与交互探索。
  • 零样本识别物体功能,模拟与真实场景下任务成功率超80%。
  • 适合开放环境中的智能体交互任务,尤其适用于模糊指令场景。

现有基于视觉-语言模型(VLM)的方法在面对模糊人类指令(如“我渴了”需识别杯子或饮料)时,难以在陌生环境中高效探索与行动。其核心挑战在于推理效率低、缺乏环境交互,导致实时任务规划与执行困难。为此,我们提出面向模糊指令的可及性感知交互决策与执行框架(AIDE),采用双流设计:多阶段推理(MSI)作为决策流,加速决策(ADM)作为执行流,实现零样本可及性分析与模糊指令解析。大量仿真与真实世界实验表明,AIDE在多样化开放世界场景中任务规划成功率超过80%,闭环连续执行准确率超过95%,运行频率达10 Hz,显著优于现有VLM方法。

原文摘要 · Abstract (English)

Enabling robots to explore and act in unfamiliar environments under ambiguous human instructions by interactively identifying task-relevant objects (e.g., identifying cups or beverages for "I'm thirsty") remains challenging for existing vision-language model (VLM)-based methods. This challenge stems from inefficient reasoning and the lack of environmental interaction, which hinder real-time task planning and execution. To address this, We propose Affordance-Aware Interactive Decision-Making and Execution for Ambiguous Instructions (AIDE), a dual-stream framework that integrates interactive exploration with vision-language reasoning, where Multi-Stage Inference (MSI) serves as the decision-making stream and Accelerated Decision-Making (ADM) as the execution stream, enabling zero-shot affordance analysis and interpretation of ambiguous instructions. Extensive experiments in simulation and real-world environments show that AIDE achieves the task planning success rate of over 80\% and more than 95\% accuracy in closed-loop continuous execution at 10 Hz, outperforming existing VLM-based methods in diverse open-world scenarios.

机器人交互模糊指令视觉语言模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。