arXiv:2412.13835cs.CL2024-12EMNLP被引 9

揭示视觉大模型在指代模糊问题上的盲区与过度自信

RAcQUEt: Unveiling the Dangers of Overlooked Referential Ambiguity in Visual LLMs

  • 构建了针对指代模糊的精细化数据集RACQUET,评估模型应对能力
  • 发现顶尖多模态模型在模糊情境下表现差且过度自信
  • 强调需避免刻板偏见,提升模型处理不确定性的能力

指代消解是有效沟通的关键。尽管人类能通过对话上下文自然解决歧义,当前语言模型是否具备类似能力尚不明确。本文提出RACQUET——一个精心设计的数据集,聚焦图像问答中的多种歧义类型。通过系列评测,揭示了先进多模态大模型在应对歧义时存在显著缺陷及过度自信问题。其中,针对社会偏见的子集RACQUET-BIAS显示:未能正确处理歧义会导致模型产生刻板化、有偏的回应。研究凸显了为模型配备稳健不确定性处理机制的紧迫性,避免依赖有害刻板印象。

原文摘要 · Abstract (English)

Ambiguity resolution is key to effective communication. While humans effortlessly address ambiguity through conversational grounding strategies, the extent to which current language models can emulate these strategies remains unclear. In this work, we examine referential ambiguity in image-based question answering by introducing RACQUET, a carefully curated dataset targeting distinct aspects of ambiguity. Through a series of evaluations, we reveal significant limitations and problems of overconfidence of state-of-the-art large multimodal language models in addressing ambiguity in their responses. The overconfidence issue becomes particularly relevant for RACQUET-BIAS, a subset designed to analyze a critical yet underexplored problem: failing to address ambiguity leads to stereotypical, socially biased responses. Our results underscore the urgency of equipping models with robust strategies to deal with uncertainty without resorting to undesirable stereotypes.

视觉语言模型指代歧义模型偏差多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。