发现长文本幻觉源于对上下文依赖加深,提出三阶段抑制框架。
Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- 通过设计上下文诱导幻觉,实现早期风险识别。
- 在多个基准上显著降低物体级幻觉率。
- 适合关注视觉语言模型可靠性研究者阅读。
大型视觉语言模型(LVLMs)近年来取得显著进展,但普遍存在幻觉问题,尤其在生成较长自由文本时更为明显,常被归因于累积不确定性。本文探讨:幻觉增多是否仅由长度引发?还是存在更深层机制?通过一系列初步实验发现,幻觉风险并非源于长度本身,而是长响应中对上下文一致性与完整性依赖增强所致。基于此,提出“诱导-检测-抑制”新框架:通过精心设计的上下文主动诱导幻觉,利用诱导样本进行早期高风险案例检测,并在实际解码过程中抑制潜在的物体级幻觉。该方法在所有基准测试中均实现一致且显著的性能提升,验证了框架有效性。强检测能力与幻觉缓解效果不仅支持了框架设计,更重新验证了关于上下文作用的核心假设。本研究旨在提供新视角,推动对LVLM长响应幻觉机制的深入探索。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have made significant progress in recent years but are also prone to hallucination issues. They exhibit more hallucinations in longer, free-form responses, often attributed to accumulated uncertainties. In this paper, we ask: Does increased hallucination result solely from length-induced errors, or is there a deeper underlying mechanism? After a series of preliminary experiments and findings, we suggest that the risk of hallucinations is not caused by length itself but by the increased reliance on context for coherence and completeness in longer responses. Building on these insights, we propose a novel "induce-detect-suppress" framework that actively induces hallucinations through deliberately designed contexts, leverages induced instances for early detection of high-risk cases, and ultimately suppresses potential object-level hallucinations during actual decoding. Our approach achieves consistent, significant improvements across all benchmarks, demonstrating its efficacy. The strong detection and improved hallucination mitigation not only validate our framework but, more importantly, re-validate our hypothesis on context. Rather than solely pursuing performance gains, this study aims to provide new insights and serves as a first step toward a deeper exploration of hallucinations in LVLMs' longer responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。