arXiv:2410.15778cs.CVcs.AI2024-10被引 56

通过调控隐空间稳定视觉特征,减少视觉语言模型幻觉。

Reducing Hallucinations in Vision-Language Models via Latent Space Steering

  • 在推理时干预隐空间,提升视觉特征稳定性。
  • 多指标测试显示幻觉率显著降低,效果优于基线。
  • 无需额外训练,可通用适配各类任务。

幻觉问题制约大型视觉语言模型(LVLMs)的实际应用。与大语言模型(LLMs)不同,LVLM中的幻觉通常源于视觉输入与文本输出之间的错位。本文研究了幻觉的内在机制,聚焦于LVLM独有的结构特点。我们发现,幻觉常由文本解码器对视觉输入的敏感性引发,这是图像编码器与文本解码器独立预训练所导致的自然现象。受此启发,我们提出视觉与文本干预(VTI),一种在推理阶段通过操控隐空间表示来增强视觉特征稳定性的新方法。作为一种无任务依赖的测试时干预技术,VTI无需额外成本即可应用于任意任务。大量实验表明,该方法能有效减少幻觉,在多个指标上优于基线,凸显了视觉特征稳定性在LVLM中的关键作用。

原文摘要 · Abstract (English)

Hallucination poses a challenge to the deployment of large vision-language models (LVLMs) in applications. Unlike in large language models (LLMs), hallucination in LVLMs often arises from misalignments between visual inputs and textual outputs. This paper investigates the underlying mechanisms of hallucination, focusing on the unique structure of LVLMs that distinguishes them from large language models (LLMs). We identify that hallucinations often arise from the sensitivity of text decoders to vision inputs, a natural phenomenon when image encoders and text decoders are pre-trained separately. Inspired by this, we introduce Visual and Textual Intervention (VTI), a novel technique designed to reduce hallucinations by steering latent space representations during inference to enhance the stability of vision features. As a task-agnostic test-time intervention, VTI can be easily applied to any problem without additional cost. Extensive experiments demonstrate that it can effectively reduce hallucinations and outperform baseline methods across multiple metrics, highlighting the critical role of vision feature stability in LVLMs.

视觉语言模型幻觉减少隐空间控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。