通过优化视觉标记减少大模型幻觉,提升图像理解准确性。
SeeMe: Mitigating Hallucinations in Large Vision-Language Models through Effective Visual Token Engineering

- 重构视觉标记,过滤干扰信息
- 在多个基准上降低幻觉率并提升一致性
- 无需训练,适配主流视觉语言模型
大型视觉语言模型(LVLMs)在图像描述和视觉问答等任务中取得了显著进展,但仍易产生与实际视觉输入不符的幻觉。现有方法多在解码阶段干预,却忽视了导致幻觉的关键来源:无关或噪声的视觉标记会误导解码过程。为此,我们提出SeeMe,一种无需训练的框架,将传统机器学习中的特征工程思想引入LVLMs。SeeMe通过三阶段标记工程过程重构视觉标记,抑制幻觉来源的同时保留有效视觉证据。在MME、POPE和AMBER三个基准上,针对四种LVLMs的实验表明,SeeMe能持续降低幻觉并提升输出一致性,为缓解LVLM幻觉提供了新视角。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have achieved remarkable progress in visual understanding tasks such as image captioning and visual question answering. However, they remain susceptible to hallucinations, generating content that is inconsistent with the actual visual input. Existing methods primarily intervene at the decoding stage, while overlooking a critical source of hallucinations: irrelevant or noisy visual tokens that mislead the decoding process. To address this issue, we propose SeeMe, a training-free framework that introduces the concept of feature engineering from traditional machine learning into LVLMs. SeeMe restructures visual tokens through a three-stage token engineering process to suppress hallucination sources while preserving informative visual evidence. Experiments on MME, POPE, and AMBER benchmarks across four LVLMs demonstrate that SeeMe consistently reduces hallucinations and improves output consistency, providing a novel perspective for mitigating hallucinations in LVLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。