arXiv:2603.10360cs.CV2026-03被引 2

通过操控视觉标记统一缓解多模态大模型幻觉问题。

One Token, Two Fates: A Unified Framework via Vision Token Manipulation Against MLLMs Hallucination

  • 以视觉标记为核心,分别增强与移除来平衡视觉与语言信息。
  • 在多个基准上使物体幻觉显著减少,平均提升POPE准确率2%。
  • 适合关注多模态推理鲁棒性的研究人员和开发者。

当前无训练方法针对多模态大模型(MLLM)幻觉采用分离策略:要么增强视觉信号,要么抑制语言惯性。然而,这两种方法存在关键权衡——单纯增强视觉常被强语言先验压制,而抑制语言则引入额外图像无关噪声。此外,简单组合效果不佳,亟需统一框架。本文聚焦核心资源:视觉标记,基于两大洞察:(1) 增强图像提供互补视觉语义;(2) 移除视觉标记(信息缺口)比扭曲图像(模态缺口)更精确地暴露幻觉倾向。据此,提出双模块框架:协同视觉校准(SVC)融合增强标记以强化视觉表征,因果表示校准(CRC)通过修剪标记生成潜空间负样本以纠正模型内在偏差。二者协同恢复视觉-语言平衡,在多个基准上使LLaVA-1.5的物体幻觉显著减少,平均提升POPE准确率2个百分点,推理延迟仅增加1.06倍。

原文摘要 · Abstract (English)

Current training-free methods tackle MLLM hallucination with separate strategies: either enhancing visual signals or suppressing text inertia. However, these separate methods are insufficient due to critical trade-offs: simply enhancing vision often fails against strong language prior, while suppressing language can introduce extra image-irrelevant noise. Moreover, we find their naive combination is also ineffective, necessitating a unified framework. We propose such a framework by focusing on the core asset: the vision token. Our design leverages two key insights: (1) augmented images offer complementary visual semantics, and (2) removing vision tokens (information-gap) isolates hallucination tendencies more precisely than distorting images (modality-gap). Based on these, our framework uses vision tokens in two distinct ways, both operating on latent representations: our Synergistic Visual Calibration (SVC) module incorporates augmented tokens to strengthen visual representations, while our Causal Representation Calibration (CRC) module uses pruned tokens to create latent-space negative samples for correcting internal model biases. By harmonizing these two roles, our framework effectively restores the vision-language balance, significantly reducing object hallucinations, improving POPE accuracy by an average of 2% absolute on LLaVA-1.5 across multiple benchmarks with only a 1.06x inference latency overhead.

多模态幻觉抑制视觉标记模型校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。