arXiv:2412.04939cs.CV2024-12AAAI

首次揭示多模态大模型的动词幻觉问题并提出针对性缓解方法。

Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models

  • 基于丰富动词知识进行微调,提升动作理解准确性。
  • 实验证明主流去幻觉方法对动词无效,动词幻觉普遍严重。
  • 适合关注视觉-语言对齐与动作理解的研究者参考。

多模态大语言模型(MLLMs)在光学字符识别、视觉问答、图像描述等任务中表现出色,但幻觉问题仍普遍存在。现有缓解方法主要针对物体/名词类概念的幻觉,而对动作相关动词概念的关注不足。本文首次系统研究了MLLMs中的动词幻觉现象,发现当前主流模型普遍存在严重的动词幻觉问题。我们评估了现有针对对象幻觉的缓解方法在动词幻觉上的效果,结果表明这些方法对动词幻觉基本无效。为此,我们提出一种基于丰富动词知识的微调方法,实验表明该方法能显著降低动词相关幻觉。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have garnered significant attention recently and demonstrate outstanding capabilities in various tasks such as OCR, VQA, captioning, $\textit{etc}$. However, hallucination remains a persistent issue. While numerous methods have been proposed to mitigate hallucinations, achieving notable improvements, these methods primarily focus on mitigating hallucinations about $\textbf{object/noun-related}$ concepts. Verb concepts, crucial for understanding human actions, have been largely overlooked. In this paper, to the best of our knowledge, we are the $\textbf{first}$ to investigate the $\textbf{verb hallucination}$ phenomenon of MLLMs from various perspectives. Our findings reveal that most state-of-the-art MLLMs suffer from severe verb hallucination. To assess the effectiveness of existing mitigation methods for object concept hallucination on verb hallucination, we evaluated these methods and found that they do not effectively address verb hallucination. To address this issue, we propose a novel rich verb knowledge-based tuning method to mitigate verb hallucination. The experiment results demonstrate that our method significantly reduces hallucinations related to verbs.

多模态幻觉检测动作理解大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。