发现大模型语言先验过弱,导致视觉问答出错。
LanP: Rethinking the Impact of Language Priors in Large Vision-Language Models
- 构建新基准LanP测试视觉语言模型的语言先验强度
- 25个模型在物体遮挡下准确率普遍低于50%
- 提醒模型需平衡语言与视觉信息依赖
大型视觉语言模型(LVLMs)在多种任务中表现优异,但存在幻觉问题,阻碍其实际应用。现有研究认为,强语言先验会压倒视觉信息导致幻觉。然而,语言先验对模型能力至关重要;若过于薄弱,模型将难以利用参数知识和指令理解能力,在视觉信息不足的复杂场景中完成任务。为此,本文提出新基准LanP,用于重新评估语言先验的影响。该基准包含170张图像和340个精心设计的问题。在25个主流LVLM上的实验表明,许多模型在物体部分遮挡时语言先验不足,准确率普遍低于0.5,包括GPT-4 Turbo在内多个模型表现不佳。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have shown impressive performance in various tasks. However, LVLMs suffer from hallucination, which hinders their adoption in the real world. Existing studies emphasized that the strong language priors of LVLMs can overpower visual information, causing hallucinations. However, the positive role of language priors is the key to a powerful LVLM. If the language priors are too weak, LVLMs will struggle to leverage rich parameter knowledge and instruction understanding abilities to complete tasks in challenging visual scenarios where visual information alone is insufficient. Therefore, we propose a benchmark called LanP to rethink the impact of Language Priors in LVLMs. It is designed to investigate how strong language priors are in current LVLMs. LanP consists of 170 images and 340 corresponding well-designed questions. Extensive experiments on 25 popular LVLMs reveal that many LVLMs' language priors are not strong enough to effectively aid question answering when objects are partially hidden. Many models, including GPT-4 Turbo, exhibit an accuracy below 0.5 in such a scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。