arXiv:2512.04597cs.CVcs.AI2025-12被引 4

提出新基准,让机器人学会在不确定时说‘我不知道’。

When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering

  • 构建新数据集,标注1636个需放弃回答的场景
  • 顶尖模型仅42.79%正确识别应放弃回答,人类达91.17%
  • 揭示当前模型依赖文字线索,缺乏真实情境判断力

具身问答(EQA)要求智能体理解语言、感知环境并在三维场景中导航以生成回答。现有基准假设每个问题都必须回答,但具身智能体应知晓何时信息不足而放弃回答。本文基于对500条人类查询的初步研究,发现32.4%的问题存在缺失或不明确的上下文。结合认知理论,归纳出五类需放弃回答的情形:行动可行性限制、指称不明确、偏好依赖、信息不可得、错误预设。我们在OpenEQA基础上,由标注者将清晰问题改写为对应模糊变体,构建新数据集AbstainEQA,包含1,636个放弃回答标注案例与1,636个原始开放问题,用于平衡评估。在该数据集上评估发现,即使最佳前沿模型的放弃回答召回率也仅为42.79%,而人类达到91.17%。此外,模型规模扩大、提示工程和推理能力提升仅带来微小改进,且微调模型易过拟合于文本线索。这些结果表明,放弃回答是具身交互中可靠性的基本前提,也是有效澄清的基础。

原文摘要 · Abstract (English)

Embodied Question Answering (EQA) requires an agent to interpret language, perceive its environment, and navigate within 3D scenes to produce responses. Existing EQA benchmarks assume that every question must be answered, but embodied agents should know when they do not have sufficient information to answer. In this work, we focus on a minimal requirement for EQA agents, abstention: knowing when to withhold an answer. From an initial study of 500 human queries, we find that 32.4% contain missing or underspecified context. Drawing on this initial study and cognitive theories of human communication errors, we derive five representative categories requiring abstention: actionability limitation, referential underspecification, preference dependence, information unavailability, and false presupposition. We augment OpenEQA by having annotators transform well-posed questions into ambiguous variants outlined by these categories. The resulting dataset, AbstainEQA, comprises 1,636 annotated abstention cases paired with 1,636 original OpenEQA instances for balanced evaluation. Evaluating on AbstainEQA, we find that even the best frontier model only attains 42.79% abstention recall, while humans achieve 91.17%. We also find that scaling, prompting, and reasoning only yield marginal gains, and that fine-tuned models overfit to textual cues. Together, these results position abstention as a fundamental prerequisite for reliable interaction in embodied settings and as a necessary basis for effective clarification.

具身问答放弃回答人工智能可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。