arXiv:2606.24112cs.AI2026-06

构建多语言多图谣言检测框架,实现低成本高精度的智能验证。

ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection

论文配图:ReMMD: Realistic Multilingual Multi-Image Agentic Verification for Multimodal Misinformation Detection
图 1 · 摘自论文原文
  • 设计可持久记忆的智能体,拆解多图多语种信息并构建证据链。
  • 在500条真实样本上达到41.8%准确率和39.12%宏F1值。
  • 适合研究多模态谣言检测与自动化验证的学者与工程师。

多模态谣言检测日益重要,因病毒式传播内容常融合长篇多语言叙述、多张图片、混合来源及细微跨模态误导。现有基准与方法难以匹配此场景:通常仅处理短文本、单图、二分类或单一篡改类型,而真实证据搜索下的智能体验证成本高昂。我们提出ReMMD,一个面向多模态谣言检测的真实世界多语言多图智能体验证框架。ReMMD包含ReMMDBench——一个含500个样本、2,756张图像、五种单语评估、两种跨语言设置、三种文本长度层级、多图帖子、五类真伪标签、八类失真标签、证据来源与推理说明的真实世界基准。同时包含ReMMD-Agent,一种具备持续记忆的验证器,能将帖子分解为原子点,构建可复用证据集,并输出结构化真伪判断、细粒度失真诊断与解释性推理。在自有系统、开源LVLM、MMD-Agent与T²-Agent中,ReMMD-Agent使用GPT-5.2实现最佳五类真伪识别性能(41.80%准确率,39.12%宏F1),相较MMD-Agent降低17.5%成本,相较T²-Agent降低79.9%成本。项目地址:https://dang-ai.github.io/ReMMD。

原文摘要 · Abstract (English)

Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed provenance, and subtle cross-modal framing errors. Existing benchmarks and methods remain poorly matched to this setting: they usually isolate short captions, single images, binary labels, or one manipulation source, while agentic verification remains costly under realistic evidence search. We present ReMMD, a realistic multilingual multi-image agentic verification framework for multimodal misinformation detection. ReMMD includes ReMMDBench, a real-world multimodal misinformation detection benchmark with 500 samples, 2,756 images, five monolingual evaluations, two cross-lingual settings, three text-length tiers, multi-image posts, five-way veracity labels, eight distortion labels, evidence provenance, and rationales. It also includes ReMMD-Agent, a persistent-memory verifier that decomposes posts into atomic points, builds a reusable evidence set, and predicts structured veracity verdicts, fine-grained distortion diagnoses, and explanatory rationales. Across proprietary systems, open LVLMs, MMD-Agent, and T$^2$-Agent, ReMMD-Agent obtains the best five-way veracity performance, with 41.80% accuracy and 39.12% macro-F1 using GPT-5.2, while reducing cost by 17.5% relative to MMD-Agent and 79.9% relative to T$^2$-Agent. The project is available at https://dang-ai.github.io/ReMMD.

多模态谣言检测智能体验证多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。