DEFAME用多模态专家动态验证图文真假,性能超现有方法。
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
- 分六步动态选工具搜证据,支持图文联合验证
- 在VERITE等三基准上超越所有旧方法,达新SOTA
- 适合需要实时、可解释图文事实核查的场景
虚假信息泛滥亟需可靠且可扩展的事实核查方案。我们提出动态证据驱动的多模态专家事实核查系统(DEFAME),一种模块化、零样本的多模态大模型流水线,用于开放域文本-图像声明验证。DEFAME采用六阶段流程,动态选择工具与搜索深度,提取并评估文本与视觉证据。不同于以往仅依赖文本、缺乏可解释性或完全依赖参数化知识的方法,DEFAME实现端到端验证,同时考虑声明与证据中的图像,并生成结构化多模态报告。在VERITE、AVerITeC和MOCHEG三个主流基准上的评估显示,DEFAME超越所有先前方法,确立了统一与多模态事实核查的新基准。此外,我们引入新的多模态基准ClaimReview2024+,包含超过GPT-4o知识截止时间的声明,避免数据泄露。在此基准上,DEFAME显著优于GPT-4o基线,展现出时间泛化能力与实时核查潜力。
原文摘要 · Abstract (English)
The proliferation of disinformation demands reliable and scalable fact-checking solutions. We present Dynamic Evidence-based FAct-checking with Multimodal Experts (DEFAME), a modular, zero-shot MLLM pipeline for open-domain, text-image claim verification. DEFAME operates in a six-stage process, dynamically selecting the tools and search depth to extract and evaluate textual and visual evidence. Unlike prior approaches that are text-only, lack explainability, or rely solely on parametric knowledge, DEFAME performs end-to-end verification, accounting for images in claims and evidence while generating structured, multimodal reports. Evaluation on the popular benchmarks VERITE, AVerITeC, and MOCHEG shows that DEFAME surpasses all previous methods, establishing itself as the new state-of-the-art fact-checking system for uni- and multimodal fact-checking. Moreover, we introduce a new multimodal benchmark, ClaimReview2024+, featuring claims after the knowledge cutoff of GPT-4o, avoiding data leakage. Here, DEFAME drastically outperforms the GPT-4o baselines, showing temporal generalizability and the potential for real-time fact-checking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。