arXiv:2601.16527cs.LGcs.AI2026-01被引 1

提出SARE方法,让多模态大模型彻底消除幻觉且不易反弹。

Beyond Superficial Unlearning: Sharpness-Aware Robust Erasure of Hallucinations in Multimodal LLMs

  • 将去学习建模为最小最大优化,主动平坦化损失曲面
  • 在重训练后仍保持幻觉抑制,效果优于现有方法
  • 适合需可靠生成的医疗、安防等高风险场景

多模态大模型虽强大,但易产生描述不存在实体的幻觉,影响可靠性。现有去学习方法存在结构性脆弱问题:标准擦除仅表面抑制,模型陷入尖锐极小值,轻量重训练后幻觉会灾难性反弹。为此,我们提出SARE框架,将去学习转化为目标最小最大优化,并引入目标-自适应梯度(Targeted-SAM)机制,显式平坦化幻觉概念周围的损失曲面。通过在模拟最坏参数扰动下抑制幻觉,确保去除过程对权重变化具有鲁棒性。大量实验表明,SARE在擦除效果上显著超越基线方法,同时保持良好生成质量;关键在于其在重训练与参数更新后仍能持续抑制幻觉,验证了几何稳定性的有效性。

原文摘要 · Abstract (English)

Multimodal LLMs are powerful but prone to object hallucinations, which describe non-existent entities and harm reliability. While recent unlearning methods attempt to mitigate this, we identify a critical flaw: structural fragility. We empirically demonstrate that standard erasure achieves only superficial suppression, trapping the model in sharp minima where hallucinations catastrophically resurge after lightweight relearning. To ensure geometric stability, we propose SARE, which casts unlearning as a targeted min-max optimization problem and uses a Targeted-SAM mechanism to explicitly flatten the loss landscape around hallucinated concepts. By suppressing hallucinations under simulated worst-case parameter perturbations, our framework ensures robust removal stable against weight shifts. Extensive experiments demonstrate that SARE significantly outperforms baselines in erasure efficacy while preserving general generation quality. Crucially, it maintains persistent hallucination suppression against relearning and parameter updates, validating the effectiveness of geometric stabilization.

多模态幻觉消除去学习鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。