arXiv:2605.10357cs.MMcs.AI2026-05被引 1

构建可审计的图文事实核查数据集,提升模型对真实网络信息的可信验证能力。

RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild

  • 基于真实社交媒体帖子构建图文对齐基准,标注推理路径与证据来源。
  • 证据约束下模型准确率与可信度显著提升,现有模型仍难忠实引用证据。
  • 适合研究多模态事实核查、可信AI验证的学者与开发者使用。

多模态虚假信息日益利用视觉说服力,通过篡改或重用图片强化误导性文本。本文提出 extbf{RW-Post},一个面向真实世界多模态事实核查的后置对齐基准,具备可审计标注:每个样本关联原始社交媒体帖子、推理轨迹及从人工事实核查文章中通过LLM辅助提取并审核的明确证据项。该数据集支持封闭书、证据受限和开放网络三种评估范式,可系统诊断模型的视觉定位与证据利用能力。我们提供 extbf{AgentFact} 作为参考验证基线,并在统一协议下评估强开源多模态大模型表现。实验表明存在显著提升空间:当前模型在忠实引用证据方面表现不佳,而证据受限评估能同时提升准确率与可信度。代码与数据集将公开于 https://github.com/xudanni0927/AgentFact。

原文摘要 · Abstract (English)

Multimodal misinformation increasingly leverages visual persuasion, where repurposed or manipulated images strengthen misleading text. We introduce \textbf{RW-Post}, a post-aligned \textbf{text--image benchmark} for real-world multimodal fact-checking with \emph{auditable} annotations: each instance links the original social-media post with reasoning traces and explicitly linked evidence items derived from human fact-check articles via an LLM-assisted extraction-and-auditing pipeline. RW-Post supports controlled evaluation across closed-book, evidence-bounded, and open-web regimes, enabling systematic diagnosis of visual grounding and evidence utilization. We provide \textbf{AgentFact} as a reference verification baseline and benchmark strong open-source LVLMs under unified protocols. Experiments show substantial headroom: current models struggle with faithful evidence grounding, while evidence-bounded evaluation improves both accuracy and faithfulness. Code and dataset will be released at https://github.com/xudanni0927/AgentFact.

多模态事实核查可信AILLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。