arXiv:2505.24007cs.CV2025-05被引 4

通过智能筛选输入图像,减少多模态模型幻觉,无需改动模型本身。

Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model

  • 根据问题类型动态选择去噪、边缘增强或原始图像作为输入。
  • 在HaloQuest数据集上使幻觉率降低44.3%,提升事实准确性。
  • 适合需要高可信度输出的医疗、法律等真实场景应用。

大型语言模型在多模态任务中常出现视觉幻觉,即生成与视觉输入不符的内容,严重影响其可靠性。现有研究多聚焦于事后修正或模型微调,对输入阶段的预处理探索不足。本文提出一种基于集成的自适应预处理框架,根据问题类型动态选择去噪(NR)、边缘增强(EE)或原始输入(org),无需修改模型架构或训练流程。在专为测试复杂视觉推理设计的HaloQuest数据集上,该方法通过SelfCheckGPT的自然语言推理评分,实现44.3%的幻觉率下降,证明仅通过智能输入调节即可显著提升模型输出的事实一致性。研究强调了自适应预处理在缓解幻觉中的关键作用,为构建更可靠的多模态系统提供新路径。

原文摘要 · Abstract (English)

Visual hallucinations in Large Language Models (LLMs), where the model generates responses that are inconsistent with the visual input, pose a significant challenge to their reliability, particularly in contexts where precise and trustworthy outputs are critical. Current research largely emphasizes post-hoc correction or model-specific fine-tuning strategies, with limited exploration of preprocessing techniques to address hallucination issues at the input stage. This study presents a novel ensemble-based preprocessing framework that adaptively selects the most appropriate filtering approach -- noise reduced (NR), edge enhanced (EE), or unaltered input (org) based on the type of question posed, resulting into reduced hallucination without requiring any modifications to the underlying model architecture or training pipeline. Evaluated on the `HaloQuest' dataset -- a benchmark designed to test multimodal reasoning on visually complex inputs, our method achieves a 44.3% reduction in hallucination rates, as measured by Natural Language Inference (NLI) scores using SelfCheckGPT. This demonstrates that intelligent input conditioning alone can significantly enhance factual grounding in LLM responses. The findings highlight the importance of adaptive preprocessing techniques in mitigating hallucinations, paving the way for more reliable multimodal systems capable of addressing real-world challenges.

多模态幻觉抑制预处理输入优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。