让AI回答时必须有图可依,避免胡编乱造
DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning
- 用反事实图像训练模型,强制答案与视觉证据对齐
- 在多个基准上准确率提升,证据一致性提高40%以上
- 适合需要可信推理的医疗、法律等高风险场景
多模态大模型虽已能结合图像思考,但常出现答案与图像证据不一致的问题。为此,我们提出DeFacto框架,通过正样本、反事实和随机掩码三种训练方式,使模型更关注正确视觉证据。我们构建了DeFacto-100K数据集,自动定位问题相关区域并生成反事实图像变体;基于该数据集,使用基于GRPO的强化学习训练模型,并设计三类奖励:正确作答、结构化推理、证据一致性。此外,我们还推出人工标注的DeFacto-1.5K基准,系统评估证据一致性。实验表明,DeFacto在多个基准上显著优于现有方法,在答案准确性和证据一致性方面均有大幅提升。
原文摘要 · Abstract (English)
Recent advances in multimodal language models (MLLMs) have made thinking with images a dominant paradigm for multimodal reasoning. However, existing methods still fail to ensure evidence-answer consistency, where correct answers must be supported by correct visual evidence. To address this issue, we propose DeFacto, a counterfactual reasoning framework that explicitly aligns visual evidence with final answers. Our approach integrates three complementary training paradigms: positive, counterfactual, and random-masking. We further develop a language-guided evidence construction pipeline that automatically localizes question-relevant regions and generates counterfactual variants, resulting in DeFacto-100K. Building on this dataset, we train MLLMs with GRPO-based reinforcement learning and design three complementary rewards to promote correct answering, structured reasoning, and consistent evidence selection. Moreover, we introduce DeFacto-1.5K, a human-annotated benchmark for systematically evaluating evidence-grounded consistency beyond answer accuracy. Experiments on diverse benchmarks demonstrate that DeFacto substantially improves both answer accuracy and evidence-answer consistency over strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。