通过自下而上的推理框架,减少多模态大模型的幻觉问题。
Combating Multimodal LLM Hallucination via Bottom-Up Holistic Reasoning
- 从感知与认知双层验证输入信息,融合视觉与常识知识。
- 在多个幻觉评测集上显著提升模型可靠性,有效抑制错误输出。
- 适合关注多模态模型可信性与推理鲁棒性的研究者。
多模态大语言模型(MLLMs)在视觉-语言任务中展现出前所未有的能力,但其普遍存在幻觉问题,即输出与输入数据不符。现有方法主要聚焦于感知层面的错误修复,却忽视了需要事实性常识的认知层面问题。同时,视觉输入表示方式不足也是引发视觉幻觉的关键瓶颈。此外,文本输入中的错误常导致模型误导,但这一问题长期被忽略。受人类直觉启发,本文提出一种自下而上的推理框架,系统性地通过整合感知级信息与认知级常识知识,验证并融合输入内容,从而生成更可靠的输出。大量实验表明,集成该框架后,MLLMs 在多个幻觉基准测试中表现显著提升。深入分析揭示该方法在缓解感知与认知层面幻觉方面具有巨大潜力。
原文摘要 · Abstract (English)
Recent advancements in multimodal large language models (MLLMs) have shown unprecedented capabilities in advancing various vision-language tasks. However, MLLMs face significant challenges with hallucinations, and misleading outputs that do not align with the input data. While existing efforts are paid to combat MLLM hallucinations, several pivotal challenges are still unsolved. First, while current approaches aggressively focus on addressing errors at the perception level, another important type at the cognition level requiring factual commonsense can be overlooked. In addition, existing methods might fall short in finding a more effective way to represent visual input, which is yet a key bottleneck that triggers visual hallucinations. Moreover, MLLMs can frequently be misled by faulty textual inputs and cause hallucinations, while unfortunately, this type of issue has long been overlooked by existing studies. Inspired by human intuition in handling hallucinations, this paper introduces a novel bottom-up reasoning framework. Our framework systematically addresses potential issues in both visual and textual inputs by verifying and integrating perception-level information with cognition-level commonsense knowledge, ensuring more reliable outputs. Extensive experiments demonstrate significant improvements in multiple hallucination benchmarks after integrating MLLMs with the proposed framework. In-depth analyses reveal the great potential of our methods in addressing perception- and cognition-level hallucinations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。