arXiv:2602.01769cs.LGcs.AI2026-02

用内部隐式奖励筛选生成对,有效减少多模态大模型幻觉。

IRIS: Implicit Reward-Guided Internal Sifting for Mitigating Multimodal Hallucination

  • 基于自生成偏好对与隐式奖励进行在线优化
  • 仅用5700样本即在关键基准上表现优异
  • 适合关注模型幻觉与高效对齐的研究者

多模态大语言模型的幻觉问题仍是核心挑战。尽管直接偏好优化(DPO)是关键对齐框架,现有方法常依赖昂贵的外部评估器进行打分或重写,导致离策略学习差距和离散化损失。由于无法访问内部状态,此类反馈忽略了生成过程中多模态间的细粒度冲突。为此,我们提出IRIS(隐式奖励引导的内部筛选),利用原生对数概率空间中的连续隐式奖励,保留完整信息密度并捕捉多模态内部竞争。该在线策略范式通过自生成偏好对消除学习差距。基于多模态隐式奖励筛选这些配对,确保优化信号直接解决模态冲突。大量实验表明,IRIS仅使用5700个样本即可在关键幻觉基准上达到高度竞争力的表现,且在偏好对齐过程中无需任何外部反馈。结果证实,IRIS为缓解多模态大模型幻觉提供了一种高效且原则化的范式。

原文摘要 · Abstract (English)

Hallucination remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While Direct Preference Optimization (DPO) is a key alignment framework, existing approaches often rely heavily on costly external evaluators for scoring or rewriting, incurring off-policy learnability gaps and discretization loss. Due to the lack of access to internal states, such feedback overlooks the fine-grained conflicts between different modalities that lead to hallucinations during generation. To address this issue, we propose IRIS (Implicit Reward-Guided Internal Sifting), which leverages continuous implicit rewards in the native log-probability space to preserve full information density and capture internal modal competition. This on-policy paradigm eliminates learnability gaps by utilizing self-generated preference pairs. By sifting these pairs based on multimodal implicit rewards, IRIS ensures that optimization is driven by signals that directly resolve modal conflicts. Extensive experiments demonstrate that IRIS achieves highly competitive performance on key hallucination benchmarks using only 5.7k samples, without requiring any external feedback during preference alignment. These results confirm that IRIS provides an efficient and principled paradigm for mitigating MLLM hallucinations.

多模态幻觉抑制对齐优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。