arXiv:2501.02699cs.CVcs.AI2025-01被引 9

提升视觉模型对图像的真实感知,减少大模型幻觉。

EAGLE: Enhanced Visual Grounding Minimizes Hallucinations in Instructional Multimodal Models

  • 重构对比预训练任务,增强视觉编码器的图像理解能力。
  • 在多个基准上显著降低幻觉率,无需重新训练语言部分。
  • 兼容现有模型架构,适合作为通用视觉优化方案。

大型语言模型与视觉变换器在零样本任务中表现出色,其融合形成的多模态架构具备强大的指令执行能力。然而,尽管经过大量图文预训练,这些模型仍常生成与图像事实不符的错误回答,即幻觉现象。现有缓解方法多聚焦于语言模块正则化、融合模块改进或集成多个视觉编码器以增强视觉表征。本文提出一种新方法 EAGLE,直接强化视觉组件的能力。EAGLE 完全独立于 LLM 和融合模块,作为后预训练策略,提升视觉编码器的图像定位与语言对齐能力。通过简单重构原始对比预训练任务,获得性能更优的视觉编码器,可无缝融入多模态系统,无需额外指令训练。实验表明,EAGLE 在多个挑战性基准和任务上均显著减少幻觉,效果显著。

原文摘要 · Abstract (English)

Large language models and vision transformers have demonstrated impressive zero-shot capabilities, enabling significant transferability in downstream tasks. The fusion of these models has resulted in multi-modal architectures with enhanced instructional capabilities. Despite incorporating vast image and language pre-training, these multi-modal architectures often generate responses that deviate from the ground truth in the image data. These failure cases are known as hallucinations. Current methods for mitigating hallucinations generally focus on regularizing the language component, improving the fusion module, or ensembling multiple visual encoders to improve visual representation. In this paper, we address the hallucination issue by directly enhancing the capabilities of the visual component. Our approach, named EAGLE, is fully agnostic to the LLM or fusion module and works as a post-pretraining approach that improves the grounding and language alignment of the visual encoder. We show that a straightforward reformulation of the original contrastive pre-training task results in an improved visual encoder that can be incorporated into the instructional multi-modal architecture without additional instructional training. As a result, EAGLE achieves a significant reduction in hallucinations across multiple challenging benchmarks and tasks.

多模态幻觉抑制视觉对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。