arXiv:2409.01151cs.CVcs.LG2024-09被引 2

提出无参对齐度量,揭示图像表征与幻觉的关联

Understanding Multimodal Hallucination with Parameter-Free Representation Alignment

  • 设计无训练参数的表征对齐度量Pfram,仅评估图像表征差异
  • 发现表征与真实标注对齐度越低,物体幻觉越严重
  • 适用于分析模型组件、指令影响及视觉编码器改进

幻觉是多模态大语言模型(MLLMs)中的常见问题,但其内在机制仍不明确。本文研究了哪些组件导致物体幻觉,并提出一种无需额外训练参数的无参表征对齐度量(Pfram),可衡量任意两个表征系统间的相似性,且完全避免其他因素干扰。特别地,Pfram能评估神经表征与人类表征(以图像真实标注为代理)之间的对齐程度。通过评估与物体标注的对齐情况,我们发现该度量在多种主流MLLMs中均与物体幻觉呈现强而一致的相关性,涵盖不同架构和规模的模型。此外,利用该度量,我们还探讨了不同模块的作用、文本指令的影响,以及使用替代视觉编码器等潜在改进方案。代码已开源:https://github.com/yellow-binary-tree/Pfram。

原文摘要 · Abstract (English)

Hallucination is a common issue in Multimodal Large Language Models (MLLMs), yet the underlying principles remain poorly understood. In this paper, we investigate which components of MLLMs contribute to object hallucinations. To analyze image representations while completely avoiding the influence of all other factors other than the image representation itself, we propose a parametric-free representation alignment metric (Pfram) that can measure the similarities between any two representation systems without requiring additional training parameters. Notably, Pfram can also assess the alignment of a neural representation system with the human representation system, represented by ground-truth annotations of images. By evaluating the alignment with object annotations, we demonstrate that this metric shows strong and consistent correlations with object hallucination across a wide range of state-of-the-art MLLMs, spanning various model architectures and sizes. Furthermore, using this metric, we explore other key issues related to image representations in MLLMs, such as the role of different modules, the impact of textual instructions, and potential improvements including the use of alternative visual encoders. Our code is available at: https://github.com/yellow-binary-tree/Pfram.

多模态幻觉检测表征对齐无参方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。