通过正负路径对比,让多模态模型生成更符合视觉事实的答案。
Breaking the Illusion: When Positive Meets Negative in Multimodal Decoding

- 在解码时引入正负双路径,强化视觉证据并惩罚语言先验主导的错误生成
- 在POPE、MME和CHAIR数据集上达到当前最佳性能,无需重新训练
- 适用于需要高视觉真实性的多模态问答场景,如医疗影像分析
视觉-语言模型常因过度依赖语言先验而产生物体幻觉,生成与视觉事实矛盾的内容。我们提出无需训练的推理框架正负解码(PND),直接干预解码过程以保证视觉一致性。PND源于发现模型中存在注意力失衡——视觉特征被低估。该框架设计双路径对比:正路径增强视觉证据,负路径构建反事实内容以惩罚先验主导的生成。通过对比两条路径的输出,引导生成结果更贴近真实视觉信息。在POPE、MME和CHAIR数据集上的实验表明,PND在不重训练的情况下达到当前最优表现。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) are frequently undermined by object hallucination, generating content that contradicts visual reality, due to an over-reliance on linguistic priors. We introduce Positive-and-Negative Decoding (PND), a training-free inference framework that intervenes directly in the decoding process to enforce visual fidelity. PND is motivated by our finding of an attention imbalance in VLMs, where visual features are under-weighted. Our framework introduces a dual-path contrast: a positive path that amplifies visual evidence and a negative path that constructs counterfactuals to penalize prior-dominant generation. By contrasting outputs from both paths during decoding, PND steers generation toward visually grounded results. Experiments on POPE, MME, and CHAIR demonstrate state-of-the-art performance without retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。