arXiv:2604.06714cs.AIcs.CL2026-04被引 5

区分明显与隐晦幻觉,实现对多模态模型幻觉可验证性的精细调控。

Steering the Verifiability of Multimodal AI Hallucinations

  • 基于人类反馈构建数据集,区分幻觉的可验证性类型。
  • 提出激活空间干预方法,分别针对显性和隐性幻觉设计调控探针。
  • 支持灵活调节不同场景下的幻觉可验证性需求,适用于高安全或高可用场景。

由多模态大语言模型(MLLMs)驱动的AI应用容易产生幻觉,对用户构成显著风险。关键在于,这些幻觉并非同等严重:部分幻觉内容可被用户轻易察觉(即明显幻觉),而另一些则难以发现或需更高验证成本(即隐晦幻觉)。这表明多模态AI幻觉在可验证性上存在显著差异。然而,目前尚缺乏对这一属性的可控研究。为此,我们基于4,470条人类对AI生成幻觉的反馈构建了数据集,并依据人类可验证性将幻觉分为明显和隐晦两类。进一步,我们提出一种激活空间干预方法,为两类幻觉分别学习独立探针。实验揭示,明显与隐晦幻觉引发不同的干预探针,从而实现对模型可验证性的细粒度控制。实证结果表明该方法有效,针对性干预在调节对应可验证性方面表现更优;且简单混合两类干预即可灵活适应不同场景的可验证性需求。

原文摘要 · Abstract (English)

AI applications driven by multimodal large language models (MLLMs) are prone to hallucinations and pose considerable risks to human users. Crucially, such hallucinations are not equally problematic: some hallucination contents could be detected by human users(i.e., obvious hallucinations), while others are often missed or require more verification effort(i.e., elusive hallucinations). This indicates that multimodal AI hallucinations vary significantly in their verifiability. Yet, little research has explored how to control this property for AI applications with diverse security and usability demands. To address this gap, we construct a dataset from 4,470 human responses to AI-generated hallucinations and categorize these hallucinations into obvious and elusive types based on their verifiability by human users. Further, we propose an activation-space intervention method that learns separate probes for obvious and elusive hallucinations. We reveal that obvious and elusive hallucinations elicit different intervention probes, allowing for fine-grained control over the model's verifiability. Empirical results demonstrate the efficacy of this approach and show that targeted interventions yield superior performance in regulating corresponding verifiability. Moreover, simply mixing these interventions enables flexible control over the verifiability required for different scenarios.

幻觉控制多模态可验证性干预方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。