通过重分配注意力提升视觉模型抗误导能力
Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
- 不需训练,将注意力从误导文本转向视觉关键区域
- 在多个模型上使误导率下降48.2%以上
- 适合关注大模型可信性的研究人员与开发者
大型多模态模型(LMMs)在多种任务中表现出色,但其易受用户‘煤气灯效应’——即故意输入误导或矛盾信息——的影响,严重威胁其实用可靠性。本文针对基于否定的煤气灯攻击问题,提出GasEraser:一种无需训练的方法,通过将注意力权重从误导性文本词转移到语义显著的视觉区域,抑制‘注意力黑洞’词的影响,强化对视觉线索的聚焦。实验表明,该方法在GaslightingBench数据集上对多个主流开源LMM有效。以LLaVA-v1.5-7B为例,其误导率降低48.2%,显著提升了模型鲁棒性。
原文摘要 · Abstract (English)
Large Multimodal Models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks. However, their vulnerability to user gaslighting-the deliberate use of misleading or contradictory inputs-raises critical concerns about their reliability in real-world applications. In this paper, we address the novel and challenging issue of mitigating the negative impact of negation-based gaslighting on LMMs, where deceptive user statements lead to significant drops in model accuracy. Specifically, we introduce GasEraser, a training-free approach that reallocates attention weights from misleading textual tokens to semantically salient visual regions. By suppressing the influence of "attention sink" tokens and enhancing focus on visually grounded cues, GasEraser significantly improves LMM robustness without requiring retraining or additional supervision. Extensive experimental results demonstrate that GasEraser is effective across several leading open-source LMMs on the GaslightingBench. Notably, for LLaVA-v1.5-7B, GasEraser reduces the misguidance rate by 48.2%, demonstrating its potential for more trustworthy LMMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。