arXiv:2510.17590cs.AIcs.CL2025-10

MERIT通过模块化设计提升多模态假信息检测,效果优于现有方法。

MERIT: Modular Framework for Multimodal Misinformation Detection with Web-Grounded Reasoning

论文配图:MERIT: Modular Framework for Multimodal Misinformation Detection with Web-Grounded Reasoning
图 1 · 摘自论文原文
  • 拆解验证流程为四个专用模块,分别处理视觉伪造、跨模态对齐等任务。
  • 在MMFakeBench上达81.65% F1,比GPT-4V高7.65个百分点,视觉伪造识别提升18.0。
  • 支持任意视觉语言模型,输出带引用的推理链,适合需要可解释性的场景。

我们提出MERIT,一种基于推理时模块化的多模态假信息检测框架,将验证过程分解为四个专用模块:视觉取证、跨模态对齐、检索增强型声明验证和校准判断。在MMFakeBench数据集上,使用GPT-4o-mini的MERIT达到81.65% F1,超越所有已报告的零样本基线(包括GPT-4V与MMD-Agent的74.0% F1)。在相同模型条件下对比实验表明,架构设计带来显著提升:相较于MMD-Agent,MERIT在假信息召回率上高出6.14点,其中视觉失真类提升+18.0,文本失真类提升+5.33。消融实验显示各模块功能不重叠,移除任一模块会显著影响对应类别性能,其余类别基本不受影响。对5,000个样本的测试集评估显示,其泛化能力与验证结果相差不超过0.21 F1点。该框架兼容任意指令遵循型视觉语言模型,并生成带引用的推理链,便于人工审查。

原文摘要 · Abstract (English)

We present MERIT, an inference-time modular framework for multimodal misinformation detection that decomposes verification into four specialized modules: visual forensics, cross-modal alignment, retrieval-augmented claim verification, and calibrated judgment. On MMFakeBench, MERIT with GPT-4o-mini achieves 81.65% F1, outperforming all reported zero-shot baselines including GPT-4V with MMD-Agent (74.0% F1). A controlled same-model evaluation confirms gains stem from architectural design: MERIT achieves 6.14 points higher misinformation recall than MMD-Agent under identical model conditions, with per-class gains of +18.0 on visual distortion and +5.33 on textual distortion. Ablation studies reveal non-overlapping module specialization, where removing any module disproportionately degrades its target category while leaving others intact. Test set evaluation on 5,000 samples confirms generalization within 0.21 F1 points of validation results. The framework operates with any instruction-following vision-language model and produces citation-linked rationales for human review.

假信息检测多模态模块化可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。