arXiv:2604.04692cs.CLcs.AI2026-04被引 1

提出自适应视觉证据使用框架,提升多模态事实核查准确率

Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity

论文配图:Is a Picture Worth a Thousand Words? Adaptive Multimodal Fact-Checking with Visual Evidence Necessity
图 1 · 摘自论文原文
  • 双模型协作:分析器判断是否需要视觉证据,验证器据此决策
  • 在三个数据集上,融合分析器判断后准确率显著提升
  • 适合关注多模态事实核查、避免无效图像引入的研究者

自动化事实核查对构建负责任的信息生态至关重要。尽管近期研究已从文本单模态发展到多模态,但普遍假设是加入视觉证据总能提升性能。本文挑战这一假设,表明盲目使用多模态证据反而会降低准确率。为此,我们提出AMuFC框架,采用两个协同的视觉-语言模型:分析器决定视觉证据是否必要,验证器基于检索到的证据及分析器评估结果预测声明真伪。在三个数据集上的实验表明,将分析器对视觉证据必要性的判断融入验证器预测,可显著提升验证性能。代码与数据集将开源。

原文摘要 · Abstract (English)

Automated fact-checking is a crucial task that supports a responsible information ecosystem. While recent research has progressed from text-only to multimodal fact-checking, a prevailing assumption is that incorporating visual evidence universally improves performance. In this work, we challenge this assumption and show that the indiscriminate use of multimodal evidence can reduce accuracy. To address this challenge, we propose AMuFC, a multimodal fact-checking framework that employs two collaborative vision-language models with distinct roles for the adaptive use of visual evidence: an Analyzer determines whether visual evidence is necessary for claim verification, and a Verifier predicts claim veracity conditioned on both the retrieved evidence and the Analyzer's assessment. Experimental results on three datasets show that incorporating the Analyzer's assessment of visual evidence necessity into the Verifier's prediction yields substantial improvements in verification performance. We will release all code and datasets at https://github.com/ssu-humane/AMuFC.

多模态事实核查视觉推理自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。