arXiv:2602.10023cs.CL2026-02Conference of the …被引 1

融合图文证据与可解释生成,提升假说验证准确性

MEVER: Multi-Modal and Explainable Claim Verification with Graph-based Evidence Retrieval

  • 构建双层图文图结构,实现跨模态证据检索
  • 通过标记级与证据级融合,提升多模态判断精度
  • 支持生成可信解释,适合科研与严谨审核场景

验证声明真伪通常需结合文本与视觉证据进行联合多模态推理,例如分析图表图像及其文字说明。同时,为使推理过程透明,需生成文本解释以支撑结论。然而,现有方法大多仅关注文本推理或忽略可解释性,导致结果不准确且缺乏说服力。为此,我们提出一种新模型,实现证据检索、多模态声明验证与解释生成的联合优化。针对证据检索,构建两层多模态图结构,设计图像到文本与文本到图像的推理机制以实现跨模态检索。针对声明验证,提出标记级与证据级融合策略,整合声明与证据嵌入进行多模态判断。针对解释生成,引入融合解码器机制增强可解释性。此外,由于现有数据集多为通用领域,我们构建了面向人工智能领域的科学数据集AIChartClaim,以补充该研究社区。实验表明,所提模型在多个任务上均表现优异。

原文摘要 · Abstract (English)

Verifying the truthfulness of claims usually requires joint multi-modal reasoning over both textual and visual evidence, such as analyzing both textual caption and chart image for claim verification. In addition, to make the reasoning process transparent, a textual explanation is necessary to justify the verification result. However, most claim verification works mainly focus on the reasoning over textual evidence only or ignore the explainability, resulting in inaccurate and unconvincing verification. To address this problem, we propose a novel model that jointly achieves evidence retrieval, multi-modal claim verification, and explanation generation. For evidence retrieval, we construct a two-layer multi-modal graph for claims and evidence, where we design image-to-text and text-to-image reasoning for multi-modal retrieval. For claim verification, we propose token- and evidence-level fusion to integrate claim and evidence embeddings for multi-modal verification. For explanation generation, we introduce multi-modal Fusion-in-Decoder for explainability. Finally, since almost all the datasets are in general domain, we create a scientific dataset, AIChartClaim, in AI domain to complement claim verification community. Experiments show the strength of our model.

多模态可解释性证据检索声明验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。