arXiv:2504.17748cs.RO2025-04被引 12

让机器人通过对话澄清模糊指令,提升任务成功率

Robotic Task Ambiguity Resolution via Natural Language Interaction

  • 用视觉语言模型将语言指令与场景关联,识别任务歧义
  • 在真实机器人上将任务成功率从69.6%提升至97.1%
  • 适合需要自然语言交互的智能机器人系统开发者

语言条件策略在机器人领域日益普及,使用户能用自然语言指定任务,具备高度灵活性。尽管现有研究多聚焦于提升语言条件策略的动作预测能力,对任务描述的推理仍被忽视。模糊的任务描述常导致机器人误解,进而引发策略失败。为此,我们提出AmbResVLM,一种将语言目标与观察场景对齐并显式推理任务歧义的新方法。我们在仿真和真实世界环境中广泛评估其有效性,结果表明,该方法在任务歧义检测与解析方面优于当前最先进的基线模型。真实机器人实验显示,该模型显著提升了下游机器人策略的表现,平均成功率从69.6%提升至97.1%。相关数据、代码与训练模型已公开,可访问https://ambres.cs.uni-freiburg.de。

原文摘要 · Abstract (English)

Language-conditioned policies have recently gained substantial adoption in robotics as they allow users to specify tasks using natural language, making them highly versatile. While much research has focused on improving the action prediction of language-conditioned policies, reasoning about task descriptions has been largely overlooked. Ambiguous task descriptions often lead to downstream policy failures due to misinterpretation by the robotic agent. To address this challenge, we introduce AmbResVLM, a novel method that grounds language goals in the observed scene and explicitly reasons about task ambiguity. We extensively evaluate its effectiveness in both simulated and real-world domains, demonstrating superior task ambiguity detection and resolution compared to recent state-of-the-art baselines. Finally, real robot experiments show that our model improves the performance of downstream robot policies, increasing the average success rate from 69.6% to 97.1%. We make the data, code, and trained models publicly available at https://ambres.cs.uni-freiburg.de.

机器人自然语言任务解析视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。