通过求解薛定谔桥问题,减少多模态大模型幻觉。
SchröMind: Mitigating Hallucinations in Multimodal Large Language Models via Solving the Schrödinger Bridge Problem
- 构建虚假与真实激活间的轻量级词元映射,最小化迁移成本。
- 在POPE和MME上达到当前最优性能,计算开销极小。
- 适合需要高可信度输出的医疗等关键领域应用。
多模态大语言模型(MLLMs)在多个领域取得显著进展,但在医疗等高风险场景中仍受限于持续存在的幻觉问题——生成文本与视觉输入矛盾或忽略。我们认为MLLMs能理解图像,但难以生成准确的词元序列。微小扰动可使注意力从真实状态转向虚假状态,且文本生成的自回归特性常导致错误无法纠正。为此,我们提出SchröMind——一种通过求解薛定谔桥问题减少幻觉的新框架。该方法在不破坏模型原有能力的前提下,以轻量训练建立虚假与真实激活间的词元级映射,实现最小运输成本。在POPE和MME基准上的大量实验表明,SchröMind表现优于现有方法,同时仅引入极小计算开销。
原文摘要 · Abstract (English)
Recent advancements in Multimodal Large Language Models (MLLMs) have achieved significant success across various domains. However, their use in high-stakes fields like healthcare remains limited due to persistent hallucinations, where generated text contradicts or ignores visual input. We contend that MLLMs can comprehend images but struggle to produce accurate token sequences. Minor perturbations can shift attention from truthful to untruthful states, and the autoregressive nature of text generation often prevents error correction. To address this, we propose SchröMind-a novel framework reducing hallucinations via solving the Schrödinger bridge problem. It establishes a token-level mapping between hallucinatory and truthful activations with minimal transport cost through lightweight training, while preserving the model's original capabilities. Extensive experiments on the POPE and MME benchmarks demonstrate the superiority of Schrödinger, which achieves state-of-the-art performance while introducing only minimal computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。