arXiv:2512.07687cs.CLcs.CV2025-12被引 3

通过分析模型内部表征变化,提升多模态大模型幻觉检测能力

HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs

  • 基于模型各层表征动态变化检测幻觉,突破传统依赖外部评估的局限
  • 在多个视觉语言任务上实现90%以上的幻觉识别准确率,优于现有方法
  • 适合关注多模态模型可信性与安全性研究的研究者使用

多模态大语言模型(MLLM)在视觉-语言理解任务中表现出色,但常产生与视觉内容不符的虚假描述,即幻觉现象,可能带来严重后果。当前幻觉评估主要依赖外部LLM评测器,但这些评测器自身也存在幻觉问题,且面临领域适应挑战。本文提出假设:幻觉表现为MLLM内部层间表征动态的可测量异常,不仅源于分布偏移,更体现在分层分析中特定假设的失效。基于此, extsc{HalluShift++} 将幻觉检测能力从纯文本大模型扩展至多模态场景。实验表明,该方法在多个基准数据集上达到超过90%的幻觉识别准确率,且代码已开源。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in vision-language understanding tasks. While these models often produce linguistically coherent output, they often suffer from hallucinations, generating descriptions that are factually inconsistent with the visual content, potentially leading to adverse consequences. Therefore, the assessment of hallucinations in MLLM has become increasingly crucial in the model development process. Contemporary methodologies predominantly depend on external LLM evaluators, which are themselves susceptible to hallucinations and may present challenges in terms of domain adaptation. In this study, we propose the hypothesis that hallucination manifests as measurable irregularities within the internal layer dynamics of MLLMs, not merely due to distributional shifts but also in the context of layer-wise analysis of specific assumptions. By incorporating such modifications, \textsc{\textsc{HalluShift++}} broadens the efficacy of hallucination detection from text-based large language models (LLMs) to encompass multimodal scenarios. Our codebase is available at https://github.com/C0mRD/HalluShift_Plus.

幻觉检测多模态表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。