视觉语言模型把本不该出错的幻觉当真,暴露其认知缺陷。
The Illusion-Illusion: Vision Language Models See Illusions Where There Are None
- 用反向幻觉测试模型,揭示其错误归因机制。
- 多款主流模型误将真实差异识别为幻觉,错误率超预期。
- 适合研究AI感知偏差与人类认知对比的学者参考。
幻觉不仅有趣,更是认知科学、哲学和神经科学中的诊断工具。典型幻觉体现‘实际状态’与‘呈现状态’之间的差距,有助于理解心理加工过程。幻觉也适用于检验人工智能系统,已有研究探讨计算模型是否像人一样受相同幻觉影响。本文反向使用感知幻觉,考察当前视觉语言模型的基本处理错误。通过呈现‘幻觉-幻觉’(illusion-illusions)——即本不该引发处理错误的邻近幻觉变体,如本就合理的鸭子、确实弯曲的线条、因真实尺寸不同而显得大小不一的圆等,发现多个主流视觉语言模型仍错误地将其识别为幻觉。这些失败可视为文献中已讨论的更广泛缺陷的一部分。
原文摘要 · Abstract (English)
Illusions are entertaining, but they are also a useful diagnostic tool in cognitive science, philosophy, and neuroscience. A typical illusion shows a gap between how something `really is' and how something `appears to be', and this gap helps us understand the mental processing that led to how something appears to be. Illusions are also useful for investigating artificial systems, and much research has examined whether computational models of perception fall prey to the same illusions as people. Here, I invert the standard use of perceptual illusions to examine basic processing errors in current vision language models. I present these models with illusory-illusions, neighbors of common illusions that should not elicit processing errors. These include such things as perfectly reasonable ducks, crooked lines that truly are crooked, circles that seem to have different sizes because they are, in fact, of different sizes, and so on. I show that many current vision language systems mistakenly see these illusion-illusions as illusions. I suggest that such failures are part of broader failures already discussed in the literature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。