Antidote通过合成数据提升视觉语言模型在反事实问题上的真实性,减少幻觉。
Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
- 用合成数据注入事实先验,让模型自纠正反事实错误
- 在CP-Bench上性能提升超50%,且无需外部标注或强模型监督
- 适合关注模型可靠性与真实性的研究者和开发者
大型视觉-语言模型(LVLMs)在多模态任务中表现优异,但幻觉问题依然突出,尤其是对反事实前提问题(CPQs)的回应常误信虚假前提并生成严重幻觉。本文揭示了LVLM在处理此类问题时的脆弱性,并提出统一的后训练框架Antidote,通过合成数据注入事实先验,实现模型自纠正;将缓解过程解耦为偏好优化问题。此外,构建新基准CP-Bench用于评估模型对反事实前提的处理能力。应用于LLaVA系列模型后,Antidote在CP-Bench上性能提升超过50%,POPE提升1.8–3.3%,CHAIR与SHR提升30–50%,且不依赖外部监督或人类反馈,未引入显著灾难性遗忘。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have achieved impressive results across various cross-modal tasks. However, hallucinations, i.e., the models generating counterfactual responses, remain a challenge. Though recent studies have attempted to alleviate object perception hallucinations, they focus on the models' response generation, and overlooking the task question itself. This paper discusses the vulnerability of LVLMs in solving counterfactual presupposition questions (CPQs), where the models are prone to accept the presuppositions of counterfactual objects and produce severe hallucinatory responses. To this end, we introduce "Antidote", a unified, synthetic data-driven post-training framework for mitigating both types of hallucination above. It leverages synthetic data to incorporate factual priors into questions to achieve self-correction, and decouple the mitigation process into a preference optimization problem. Furthermore, we construct "CP-Bench", a novel benchmark to evaluate LVLMs' ability to correctly handle CPQs and produce factual responses. Applied to the LLaVA series, Antidote can simultaneously enhance performance on CP-Bench by over 50%, POPE by 1.8-3.3%, and CHAIR & SHR by 30-50%, all without relying on external supervision from stronger LVLMs or human feedback and introducing noticeable catastrophic forgetting issues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。