arXiv:2505.01958cs.CVcs.CL2025-05被引 6

剖析大模型视觉幻觉成因并提出针对性缓解方案

A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models

  • 拆解视觉语言模型各组件,定位幻觉源头
  • 提出分组件缓解方法,有效降低幻觉率
  • 构建两个新评测基准,覆盖属性与认知幻觉

大型视觉语言模型(LVLMs)在多模态任务中表现出色,但视觉对象幻觉问题依然存在。该现象指模型基于查询输入生成不准确的视觉对象相关信息,可能导致误导性信息,引发安全与可靠性担忧。以往研究集中于幻觉的评估与缓解,但对根本原因缺乏全面分析。本文系统分析了类似LLaVA的LVLM架构中语言模型、视觉主干网络和投影模块的潜在错误来源及其影响。基于观察结果,针对每个问题组件提出缓解方法。此外,我们构建了两个幻觉评测基准:QA-VisualGenome,侧重属性与关系幻觉;QA-FB15k,聚焦认知类幻觉。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) demonstrate remarkable capabilities in multimodal tasks, but visual object hallucination remains a persistent issue. It refers to scenarios where models generate inaccurate visual object-related information based on the query input, potentially leading to misinformation and concerns about safety and reliability. Previous works focus on the evaluation and mitigation of visual hallucinations, but the underlying causes have not been comprehensively investigated. In this paper, we analyze each component of LLaVA-like LVLMs -- the large language model, the vision backbone, and the projector -- to identify potential sources of error and their impact. Based on our observations, we propose methods to mitigate hallucination for each problematic component. Additionally, we developed two hallucination benchmarks: QA-VisualGenome, which emphasizes attribute and relation hallucinations, and QA-FB15k, which focuses on cognition-based hallucinations.

视觉幻觉大模型多模态评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。