arXiv:2605.08245cs.CVcs.AI2026-05被引 3

发现视觉语言模型因过度对齐文本特征而产生幻觉,提出几何去偏方法有效抑制。

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

论文配图:When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models
图 1 · 摘自论文原文
  • 揭示解码型模型因几何过对齐导致视觉信息被语言主导
  • 在多个基准上减少幻觉,长文本生成准确率提升
  • 无需训练的推理方法可零额外开销部署

视觉语言模型广泛应用于医疗影像与自动驾驶等高风险场景,却常出现自信地描述输入中并不存在的内容。我们通过机制分析发现,这类错误源于解码式模型的几何过对齐:为弥合模态差距,模型将视觉嵌入过度对齐至文本流形,引入统计性语言偏差,系统性压制细粒度视觉证据。现有方法或强行闭合模态差距,或依赖昂贵黑盒解码策略抑制幻觉,均未触及根本成因。本文首次定量刻画该过对齐现象,证明语言偏差集中于一个通用、数据无关的文本子空间的前主成分。基于此,提出两种互补方案:免训练推理策略与偏差感知微调范式,均通过从视觉表示中显式投影出该子空间实现去偏。在POPE、CHAIR和AMBER等多个基准上显著降低幻觉率,并提升CLAIR长文本生成评分,其中免训练版本不增加任何计算开销。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) increasingly power high-stakes applications, from medical imaging to autonomous systems, yet they routinely hallucinate, confidently describing content not present in the input. We investigate the root causes of these failure modes with a mechanistic analysis focusing on the decoder-based VLMs. We trace these failure modes to a geometric over-alignment: to bridge the modality gap required by attention mechanisms, decoder-based VLMs over-align visual embeddings with the text manifold, injecting a statistical linguistic bias that systematically overshadows fine-grained visual evidence. While prior work either aggressively closes this gap or suppresses hallucinations through expensive black-box decoding strategies, none addresses the underlying geometric cause. We provide the first quantitative characterization of this over-alignment, demonstrating that linguistic bias concentrates in the top principal components of a universal, dataset-agnostic text subspace. Building on this insight, we propose two complementary remedies: a training-free inference strategy and a bias-aware fine-tuning paradigm, both of which explicitly project out this subspace from visual representations. Our methods significantly reduce hallucinations across POPE, CHAIR, and AMBER benchmarks, and improve CLAIR scores on long-form captioning tasks, with the training-free variant adding no computational overhead over the base model.

视觉语言模型幻觉抑制去偏几何分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。