通过视觉变化提升模型对细节的理解,减少幻觉。
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
- 生成可控变化的视觉图像,增强视觉与文本对齐。
- 在多个基准上显著降低幻觉率,提升细粒度理解能力。
- 适合需要高精度视觉推理的研究者和应用开发。
大型视觉-语言模型(LVLMs)在视觉理解与多模态推理方面展现出强大能力,但常出现与输入图像内容不一致的幻觉现象。现有方法多聚焦文本,受限于视觉-语义对齐问题,尤其在细粒度视觉理解场景下效果不佳。为此,本文提出ViHallu——一种以视觉为中心的幻觉缓解框架,通过视觉变化图像生成与视觉指令构建,增强视觉-语义对齐。ViHallu生成保持整体结构但具可控视觉差异的图像,并结合精心设计的视觉指令进行微调,使模型更精准捕捉视觉内容与文本的对应关系。大量实验表明,ViHallu有效提升了模型的细粒度视觉理解能力,显著减少幻觉倾向。此外,我们发布了专门用于幻觉缓解与视觉-语义对齐的视觉指令数据集ViHallu-Instruction。代码已开源:https://github.com/oliviadzy/ViHallu。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual understanding and multimodal reasoning. However, LVLMs frequently exhibit hallucination phenomena, manifesting as the generated textual responses that demonstrate inconsistencies with the provided visual content. Existing hallucination mitigation methods are predominantly text-centric, the challenges of visual-semantic alignment significantly limit their effectiveness, especially when confronted with fine-grained visual understanding scenarios. To this end, this paper presents ViHallu, a Vision-Centric Hallucination mitigation framework that enhances visual-semantic alignment through Visual Variation Image Generation and Visual Instruction Construction. ViHallu introduces visual variation images with controllable visual alterations while maintaining the overall image structure. These images, combined with carefully constructed visual instructions, enable LVLMs to better understand fine-grained visual content through fine-tuning, allowing models to more precisely capture the correspondence between visual content and text, thereby enhancing visual-semantic alignment. Extensive experiments on multiple benchmarks show that ViHallu effectively enhances models' fine-grained visual understanding while significantly reducing hallucination tendencies. Furthermore, we release ViHallu-Instruction, a visual instruction dataset specifically designed for hallucination mitigation and visual-semantic alignment. Code is available at https://github.com/oliviadzy/ViHallu.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。