让视觉语言模型学会因果推理,提升细节判断和真实性
CF-VLM:CounterFactual Vision-Language Fine-tuning

- 用反事实样本训练,增强模型对因果关系的感知能力
- 在组合推理任务上超越现有方法,减少视觉幻觉现象
- 适合需要可靠推理与可解释性的高风险应用场景
近年来视觉语言模型(VLMs)在跨模态语义理解方面取得显著进展,但在细粒度区分和深层因果推理任务中仍存在明显局限。现有模型常依赖表面统计相关性,缺乏捕捉视觉与文本内容间潜在因果逻辑的能力。为此,我们提出反事实视觉语言微调(CF-VLM),通过有针对性地使用反事实样本,提升VLM的因果推理能力。CF-VLM引入三个互补的训练目标:保持基础跨模态对齐、强化真实场景表征对一致反事实样本的唯一性与稳定性、以及提升模型对微小但关键因果修改的敏感度。大量实验表明,CF-VLM在组合推理与泛化基准测试中持续优于强基线及前沿方法,且在缓解视觉幻觉方面表现良好,体现出更强的事实一致性。本工作为部署需高可靠性推理与可解释性的VLMs提供了坚实基础。
原文摘要 · Abstract (English)
Recent advances in vision-language models (VLMs) have greatly improved cross-modal semantic understanding, yet significant limitations remain in fine-grained discrimination and deep causal reasoning tasks. Existing VLMs often rely on superficial statistical correlations, lacking the ability to capture the underlying causal logic between visual and textual content. To address this, we propose CounterFactual Vision-Language Fine-tuning (CF-VLM), a novel framework that enhances the causal reasoning capabilities of VLMs through the targeted use of counterfactual samples. CF-VLM introduces three complementary training objectives: maintaining foundational cross-modal alignment, reinforcing the uniqueness and stability of factual scene representations against coherent counterfactuals, and sharpening the model's sensitivity to minimal but critical causal edits. Extensive experiments demonstrate that CF-VLM consistently outperforms strong baselines and state-of-the-art methods on compositional reasoning and generalization benchmarks. Furthermore, it shows promise in mitigating visual hallucinations, indicating improved factual consistency. Our CF-VLM provides a robust foundation for deploying VLMs in high-stakes, real-world scenarios requiring reliable reasoning and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。