arXiv:2508.04201cs.CVcs.AI2025-08

提出ViFP框架,检测并修正视觉语言模型的错误推理路径。

ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs

  • 通过多轮问答构建推理链,动态分析一致性以识别假阳性推理。
  • 在A-OKVQA上提升准确率5.4%,优于此前最优方法4.3%。
  • 适合关注模型推理可靠性与可解释性的研究者使用。

在视觉语言模型(VLMs)推理过程中,当模型得出正确答案但遵循错误推理路径时,会出现假阳性(FP)推理,削弱推理可靠性。现有方法多依赖提示工程、知识蒸馏或强化学习,均需大量高质量数据,限制了实际应用。少数研究关注直接检测和纠正FP。为此,本文提出ViFP框架,通过多轮问答构建有效推理链,并动态分析推理路径一致性以识别潜在FP。同时引入针对性推理链修正机制,改善逻辑一致性和准确性。最后,提出可靠性评估指标VoC,融合答案准确率与FP率,量化评估模型是否不仅答对,且推理可靠。在闭源VLM上实验表明,ViFP在A-OKVQA、OK-VQA和FVQA三个数据集上均持续提升性能,在A-OKVQA上准确率最高提升5.4%,超越此前最先进方法4.3%,显著减少FP数量,验证其在增强推理可靠性方面的有效性。

原文摘要 · Abstract (English)

During reasoning in vision-language models (VLMs), false positive (FP) reasoning occurs when a model produces the correct answer but follows an incorrect reasoning path, resulting in undermined reasoning reliability. Existing approaches mainly rely on prompt engineering, knowledge distillation or reinforcement learning to improve reasoning reliability, both of which require large amounts of high-quality data and thus limit practical applicability. Few approaches have focused on directly detecting and correcting FPs. To address these issues, we propose ViFP, a framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs. ViFP builds effective reasoning paths through multi-turn QA and dynamically analyzes the consistency of the reasoning path to identify potential FPs. It also introduces a targeted reasoning chain correction mechanism to modify FP reasoning, thereby improving logical consistency and accuracy. Finally, we introduce a reliability evaluation metric, VoC, which integrates answer accuracy and the FP rate, providing a quantitative tool to assess whether a VLM not only answers correctly but also reasons reliably. Our experiments on closed-source VLMs show that ViFP consistently improves performance across three datasets: A-OKVQA, OK-VQA, and FVQA. On A-OKVQA, ViFP improves accuracy by up to 5.4%, surpassing the previous state-of-the-art by 4.3%, and significantly reduces the number of FPs, validating its benefits in enhancing reasoning reliability.

视觉语言模型推理可靠性假阳性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。