arXiv:2412.10726cs.CVcs.AI2024-12被引 12

评测智能体在含噪声问题下的问答能力,提升真实场景适应性。

NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries

  • 构建含四类噪声的NoisyEQA基准,模拟真实提问干扰。
  • 现有模型难识别噪声,导致错误回答率超50%。
  • 提出自纠正提示机制,显著提升答案准确率。

视觉语言模型的快速发展推动了具身问答(EQA)的进步,提升了智能体在复杂现实场景中的语言理解与推理能力。然而,在真实场景中,人类提出的疑问常包含噪声,干扰智能体探索与回应,尤其对语言初学者和非专业用户构成挑战。为此,我们提出NoisyEQA基准,用于评估智能体识别并纠正噪声问题的能力。该基准通过自动化数据生成框架引入四种常见噪声:潜在幻觉噪声、记忆噪声、感知噪声和语义噪声。此外,我们还提出一种“自纠正”提示机制与新的评估指标,以增强并量化噪声检测能力与回答质量。全面评估表明,当前EQA智能体普遍难以识别问题中的噪声,导致回答频繁包含错误信息。通过自纠正提示机制,可有效提升智能体回答准确性。

原文摘要 · Abstract (English)

The rapid advancement of Vision-Language Models (VLMs) has significantly advanced the development of Embodied Question Answering (EQA), enhancing agents' abilities in language understanding and reasoning within complex and realistic scenarios. However, EQA in real-world scenarios remains challenging, as human-posed questions often contain noise that can interfere with an agent's exploration and response, bringing challenges especially for language beginners and non-expert users. To address this, we introduce a NoisyEQA benchmark designed to evaluate an agent's ability to recognize and correct noisy questions. This benchmark introduces four common types of noise found in real-world applications: Latent Hallucination Noise, Memory Noise, Perception Noise, and Semantic Noise generated through an automated dataset creation framework. Additionally, we also propose a 'Self-Correction' prompting mechanism and a new evaluation metric to enhance and measure both noise detection capability and answer quality. Our comprehensive evaluation reveals that current EQA agents often struggle to detect noise in questions, leading to responses that frequently contain erroneous information. Through our Self-Correct Prompting mechanism, we can effectively improve the accuracy of agent answers.

具身问答噪声鲁棒视觉语言模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。