通过精准调控关键视觉敏感词元,有效减少大模型幻觉。
Steer Where It Matters: Token-Level Visual-Sensitivity Steering for LVLMs Hallucination Mitigation

- 按词元粒度提取并优化视觉敏感性引导向量
- 仅在需要时动态调整干预强度,显著提升抑制效果
- 轻量插件式设计,适配多种视觉语言模型
大型视觉语言模型(LVLMs)虽快速进步并广泛应用,但幻觉问题仍严重。激活引导因其训练开销小、推理时可控而具吸引力。然而我们发现,在自回归解码中,视觉条件仅稀疏且局部影响词元预测,现有方法对整个序列平均图像与无图像差异,稀释了关键信号,导致信噪比低。此外,多数方法采用固定引导强度,错误分配干预预算,过度扰动非关键词元,引发不稳定性。为此,我们提出词元级视觉敏感性引导(TLVS),先提取并精炼词元级引导向量,再在关键位置实施细粒度、自适应引导。该轻量、即插即用机制仅需少量校准训练,可广泛应用于不同视觉语言模型。它在每个解码步骤动态调节引导强度,选择性抑制易产生幻觉的片段,同时保留基于证据的内容。我们在POPE、AMBER、CHAIR(COCO)、MMHal和HallusionBench等多个基准上评估,结果一致优于以往引导方法。
原文摘要 · Abstract (English)
Large vision language models (LVLMs) have made rapid advancements and are deployed across various applications, yet hallucinations remain a major challenge. Activation steering is appealing due to its minimal training overhead and controllability at inference time. However, we found that during autoregressive decoding, visual conditioning affects token prediction sparsely and locally across decoding steps, and many existing methods that average image-versus-no-image differences over the entire sequence dilute these critical signals, yielding low signal-to-noise ratio steering directions. Additionally, many existing methods apply a fixed steering strength, which misallocates the intervention budget, over-perturbs non-critical tokens, and can cause instability. To address these limitations, we propose Token-Level Visual-Sensitivity Steering (TLVS) for hallucination mitigation. Our approach first extracts token-level steering vectors and refines them, and then applies fine-grained, visual-sensitivity-adaptive steering only where it matters. This lightweight, plug-and-play mechanism requires only minimal training for calibration and can be applied across diverse vision-language models. It modulates the steering strength at each decoding step, selectively suppressing hallucination-prone spans while preserving evidence-grounded content. We evaluate TLVS on several benchmarks, including POPE, AMBER, CHAIR (COCO), MMHal, and HallusionBench, demonstrating consistent improvements over previous steering methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。