arXiv:2604.00983cs.CV2026-04

通过动态整合上下文信息,有效减少大模型视觉幻觉。

ACT Now: Preempting LVLM Hallucinations via Adaptive Context Integration

  • 动态调整注意力头,增强视觉信息探索能力
  • 通过语义聚合缓解离散生成导致的信息损失
  • 无需训练,适配多种模型且不降低生成质量

大型视觉语言模型(LVLMs)普遍存在严重幻觉问题。现有方法多依赖静态、单步策略增强视觉关注或抑制语言先验,但忽视生成过程中上下文的动态变化,难以纠正信息丢失。为此,我们提出无需训练的推理干预方法ACT(自适应上下文整合),通过视觉上下文探索和语义上下文聚合来缓解幻觉。前者利用时空分析自适应增强负责视觉探索的注意力头;后者通过边缘化潜在语义查询,有效聚合视觉证据,解决离散标记预测带来的信息损失。在多种LVLM上广泛实验表明,ACT显著降低幻觉率,在判别与生成基准上均取得竞争力表现,兼具鲁棒性与高适应性,且不损害基础生成能力。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) frequently suffer from severe hallucination issues. Existing mitigation strategies predominantly rely on isolated, single-step states to enhance visual focus or suppress strong linguistic priors. However, these static approaches neglect dynamic context changes across the generation process and struggles to correct inherited information loss. To address this limitation, we propose Adaptive Context inTegration (ACT), a training-free inference intervention method that mitigates hallucination through the adaptive integration of contextual information. Specifically, we first propose visual context exploration, which leverages spatio-temporal profiling to adaptively amplify attention heads responsible for visual exploration. To further facilitate vision-language alignment, we propose semantic context aggregation that marginalizes potential semantic queries to effectively aggregate visual evidence, thereby resolving the information loss caused by the discrete nature of token prediction. Extensive experiments across diverse LVLMs demonstrate that ACT significantly reduces hallucinations and achieves competitive results on both discriminative and generative benchmarks, acting as a robust and highly adaptable solution without compromising fundamental generation capabilities.

视觉语言模型幻觉抑制上下文整合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。