arXiv:2508.03177cs.CV2025-08AAAI被引 2

针对风格化图像引发的幻觉问题,提出早期视觉修正机制。

SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision

论文配图:SAVER: Mitigating Hallucinations in Large Vision-Language Models via Style-Aware Visual Early Revision
图 1 · 摘自论文原文
  • 基于视觉注意力模式动态调整输出,实现早期修正
  • 在13个LVLM上验证,显著降低风格化图像幻觉率
  • 适合游戏、医疗等对图像风格敏感的场景

大型视觉语言模型(LVLM)在理解复杂图文上下文方面取得显著进展,但幻觉问题仍限制其实际应用。尽管现有方法能有效减少照片类图像的幻觉,却忽视了风格化图像带来的风险,而这类图像在游戏场景理解、艺术教育和医学分析中至关重要。本文构建了一个包含照片及其对应风格化版本的数据集,并对13个先进LVLM进行了判别与生成任务的对比实验。结果表明,风格化图像比照片诱发的幻觉更严重。为此,我们提出风格感知的视觉早期修正机制SAVER,通过利用早期层的视觉注意力模式进行动态输出调整,缓解由风格化图像引起的幻觉。大量实验表明,SAVER在多种模型、数据集和任务中均达到当前最优的幻觉抑制效果。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) recently achieve significant breakthroughs in understanding complex visual-textual contexts. However, hallucination issues still limit their real-world applicability. Although previous mitigation methods effectively reduce hallucinations in photographic images, they largely overlook the potential risks posed by stylized images, which play crucial roles in critical scenarios such as game scene understanding, art education, and medical analysis. In this work, we first construct a dataset comprising photographic images and their corresponding stylized versions with carefully annotated caption labels. We then conduct head-to-head comparisons on both discriminative and generative tasks by benchmarking 13 advanced LVLMs on the collected datasets. Our findings reveal that stylized images tend to induce significantly more hallucinations than their photographic counterparts. To address this issue, we propose Style-Aware Visual Early Revision SAVER, a novel mechanism that dynamically adjusts LVLMs' final outputs based on the token-level visual attention patterns, leveraging early-layer feedback to mitigate hallucinations caused by stylized images. Extensive experiments demonstrate that SAVER achieves state-of-the-art performance in hallucination mitigation across various models, datasets, and tasks.

视觉语言模型幻觉抑制风格化图像早期修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。