用双向最大似然方法提升视觉语言模型的幻觉预测与抑制能力
BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models
- 基于归一化流理论构建双向最大似然学习框架
- 在POPE上达85.06%平均F1,CHAIRS/CHAIRI降低7.6%/2.6%
- 首次将双射性用于减少大模型幻觉,适合多模态研究者
大型视觉语言模型在多个领域广泛应用,但其可解释性差导致可信系统构建困难。其中最常见问题为幻觉——语言模型生成与视觉内容不符的响应。为缓解此问题,已有研究聚焦解码过程优化。本文提出一种基于归一化流理论的双向最大似然学习(BIMA)方法,有效抑制主流视觉语言模型中的幻觉现象。实验表明,BIMA在POPE基准上实现85.06%的平均F1分数,并使CHAIRS和CHAIRI指标分别降低7.6%和2.6%。据我们所知,这是首个探讨双射性机制以减少大模型幻觉的研究。
原文摘要 · Abstract (English)
Large vision-language models have become widely adopted to advance in various domains. However, developing a trustworthy system with minimal interpretable characteristics of large-scale models presents a significant challenge. One of the most prevalent terms associated with the fallacy functions caused by these systems is hallucination, where the language model generates a response that does not correspond to the visual content. To mitigate this problem, several approaches have been developed, and one prominent direction is to ameliorate the decoding process. In this paper, we propose a new Bijective Maximum Likelihood Learning (BIMA) approach to hallucination mitigation using normalizing flow theories. The proposed BIMA method can efficiently mitigate the hallucination problem in prevailing vision-language models, resulting in significant improvements. Notably, BIMA achieves the average F1 score of 85.06% on POPE benchmark and remarkably reduce CHAIRS and CHAIRI by 7.6% and 2.6%, respectively. To the best of our knowledge, this is one of the first studies that contemplates the bijection means to reduce hallucination induced by large vision-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。