系统梳理大视觉语言模型幻觉问题及其应对策略
A Survey of Hallucination in Large Visual Language Models
- 分析视觉语言模型结构与幻觉成因
- 归纳近年幻觉修正与缓解方法
- 提供评估幻觉的判断与生成类基准
大视觉语言模型(LVLM)在大型语言模型基础上融合视觉模态,显著提升用户交互体验和信息处理能力。然而,幻觉现象严重制约了其在各领域的实际应用。尽管已有大量研究致力于幻觉缓解与修正,但相关综述仍较少。本文首先介绍LVLM与幻觉的背景,阐述其模型结构及幻觉生成的主要原因。进一步总结近期幻觉修正与缓解方法,从判断与生成两个视角梳理现有的幻觉评估基准。最后,提出未来研究方向,以增强LVLM的可靠性与实用性。
原文摘要 · Abstract (English)
The Large Visual Language Models (LVLMs) enhances user interaction and enriches user experience by integrating visual modality on the basis of the Large Language Models (LLMs). It has demonstrated their powerful information processing and generation capabilities. However, the existence of hallucinations has limited the potential and practical effectiveness of LVLM in various fields. Although lots of work has been devoted to the issue of hallucination mitigation and correction, there are few reviews to summary this issue. In this survey, we first introduce the background of LVLMs and hallucinations. Then, the structure of LVLMs and main causes of hallucination generation are introduced. Further, we summary recent works on hallucination correction and mitigation. In addition, the available hallucination evaluation benchmarks for LVLMs are presented from judgmental and generative perspectives. Finally, we suggest some future research directions to enhance the dependability and utility of LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。