揭秘视觉语言模型少幻觉的架构秘诀,给出可落地的设计指南。
What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness

- 从语言基础、视觉表征、语义对齐三维度分析幻觉成因。
- 大模型参数量影响有限,更强语言基座和视觉编码器能降幻觉。
- 视觉质量与对齐策略协同提升,效果最佳,适合模型开发者参考。
幻觉仍是制约大型视觉语言模型(LVLMs)可靠性的关键问题。我们提出,幻觉的根本原因在于模型架构设计。将架构设计分解为语言基础(LF)、视觉表征(VR)和语义对齐(SA)三个维度,并将幻觉分为共现、相似性和未被重视的不确定性三类。基于此,我们提出CoSimUE基准,通过受控文本扰动和随机扰动生成细粒度幻觉场景,实现设计选择与幻觉行为的映射。在7个设计方面的实验表明:1)模型参数量扩大对三类幻觉缓解作用有限;2)更大更优的语言基础可减少共现幻觉;3)更强的视觉编码器和更高分辨率能减轻相似性错误;4)有效的对齐策略可缓解不确定性幻觉;5)跨维度分析显示,同时提升视觉保真度和对齐质量带来最全面的改进。本研究首次系统性揭示架构设计与幻觉鲁棒性的关联,为开发可靠高效的LVLM提供实践指导。
原文摘要 · Abstract (English)
Hallucination remains one of the key challenges undermining the reliability of Large Vision-Language Models (LVLMs). But what makes an LVLM hallucinate less? Many existing efforts focus on improving internal components of the model. We argue that hallucination fundamentally stems from how the model architecture is designed. To investigate this, we factor the architecture design into three dimensions: Linguistic Foundation (LF), Visual Representation (VR), and Semantic Alignment (SA), and categorize hallucinations into Co-occurrence, Similarity, and previously overlooked Uncertainty types. Building on this formulation, we propose CoSimUE, a benchmark that creates fine-grained hallucination scenarios through controlled textual perturbations and random perturbations, enabling mapping between design choices and hallucination behaviors. Experiments across 7 design aspects show that: 1) the widely emphasized scaling of model parameters has only limited impact on reducing all three types of hallucinations; 2) larger and better-trained language foundations can reduce co-occurrence hallucinations; 3) stronger visual encoders and higher resolutions mitigate similarity errors; 4) effective alignment strategies alleviate uncertainty hallucinations. 5) Furthermore, cross-dimensional analysis reveals that jointly enhancing visual fidelity and alignment quality yields the most comprehensive improvements. This study provides the first systematic exploration linking architecture-level design to hallucination robustness, offering practical guidance for developing reliable and efficient LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。