通过自适应控制信息流,减少视觉语言模型的幻觉问题
Mitigating Hallucinations in Large Vision-Language Models by Adaptively Constraining Information Flow
- 用变分信息瓶颈引入随机噪声,缓解无关视觉特征的过度自信
- 基于相似度分布平滑性动态调节噪声强度,提升约束精度
- 在多个模型架构上验证有效,显著降低图像描述中的幻觉率
大型视觉语言模型在通过人类语言理解视觉信息方面展现出巨大潜力,但容易出现物体幻觉——生成的图像描述包含图像中并不存在的物体。本文揭示,物体幻觉源于软视觉标记映射到大语言模型词嵌入空间时对无关视觉特征的过度自信。通过分析视觉标记与语言模型词嵌入之间的语义相似度,我们发现相似度分布的平滑性与幻觉出现高度相关。为此,提出使用变分信息瓶颈(VIB)引入随机噪声,以缓解过度自信,并设计基于熵的噪声控制策略,使注入噪声能自适应地根据相似度分布平滑性进行约束。该方法在多种模型架构上均取得一致效果,在两个物体幻觉基准测试中显著降低幻觉率。
原文摘要 · Abstract (English)
Large vision-language models show tremendous potential in understanding visual information through human languages. However, they are prone to suffer from object hallucination, i.e., the generated image descriptions contain objects that do not exist in the image. In this paper, we reveal that object hallucination can be attributed to overconfidence in irrelevant visual features when soft visual tokens map to the LLM's word embedding space. Specifically, by figuring out the semantic similarity between visual tokens and LLM's word embedding, we observe that the smoothness of similarity distribution strongly correlates with the emergence of object hallucinations. To mitigate hallucinations, we propose using the Variational Information Bottleneck (VIB) to alleviate overconfidence by introducing stochastic noise, facilitating the constraining of irrelevant information. Furthermore, we propose an entropy-based noise-controlling strategy to enable the injected noise to be adaptively constrained regarding the smoothness of the similarity distribution. We adapt the proposed AdaVIB across distinct model architectures. Experimental results demonstrate that the proposed AdaVIB mitigates object hallucinations by effectively alleviating the overconfidence in irrelevant visual features, with consistent improvements on two object hallucination benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。