arXiv:2601.05991cs.AI2026-01被引 2

检测3D场景中指令的模糊性,提升机器人安全执行能力。

3D Instruction Ambiguity Detection

  • 构建多视角视觉证据,判断指令在3D场景中是否唯一
  • 22000条指令测试显示顶尖3D大模型难以可靠识别模糊指令
  • 适合关注机器人安全与人机交互可信度的研究者

在安全关键领域,语言模糊可能导致严重后果;例如手术中一句“把试管递给我”可能引发灾难性错误。然而,当前具身AI研究普遍忽视此问题,假设指令清晰而只关注执行。为此,我们首次提出3D指令模糊性检测任务:模型需判断给定3D场景中的指令是否有唯一明确含义。为此,我们构建了大规模基准Ambi3D,包含700多个多样化的3D场景和约22,000条指令。分析发现,现有最先进的3D大语言模型(LLMs)在判断指令模糊性方面表现不佳。为此,我们提出AmbiVer框架,通过多视角获取显式视觉证据,并引导视觉语言模型(VLM)进行判断。大量实验验证了该任务的挑战性及AmbiVer的有效性,为更安全、可信赖的具身AI铺平道路。代码与数据集见https://jiayuding031020.github.io/ambi3d/。

原文摘要 · Abstract (English)

In safety-critical domains, linguistic ambiguity can have severe consequences; a vague command like "Pass me the vial" in a surgical setting could lead to catastrophic errors. Yet, most embodied AI research overlooks this, assuming instructions are clear and focusing on execution rather than confirmation. To address this critical safety gap, we are the first to define 3D Instruction Ambiguity Detection, a fundamental new task where a model must determine if a command has a single, unambiguous meaning within a given 3D scene. To support this research, we build Ambi3D, the large-scale benchmark for this task, featuring over 700 diverse 3D scenes and around 22k instructions. Our analysis reveals a surprising limitation: state-of-the-art 3D Large Language Models (LLMs) struggle to reliably determine if an instruction is ambiguous. To address this challenge, we propose AmbiVer, a two-stage framework that collects explicit visual evidence from multiple views and uses it to guide an vision-language model (VLM) in judging instruction ambiguity. Extensive experiments demonstrate the challenge of our task and the effectiveness of AmbiVer, paving the way for safer and more trustworthy embodied AI. Code and dataset available at https://jiayuding031020.github.io/ambi3d/.

具身智能指令理解安全检测3D感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。