用多智能体检索增强框架,自动验证图文谣言真伪。
RAMA: Retrieval-Augmented Multi-Agent Framework for Misinformation Detection in Multimodal Fact-Checking
- 将模糊图文内容转为精准网络搜索词,定位可信证据。
- 融合多方权威来源信息,提升对可疑内容的判别力。
- 适合需要高可靠性的媒体审核与社交平台风控场景。
多模态谣言的快速传播给自动化事实核查系统带来严峻挑战,尤其当声明模糊或缺乏上下文时。我们提出RAMA,一种新型检索增强型多智能体框架,用于验证多媒体谣言。RAMA包含三项核心创新:(1)策略性查询生成,将多模态声明转化为精确的网络搜索查询;(2)来自多样且权威来源的交叉验证证据聚合;(3)多智能体集成架构,利用多个多模态大语言模型及提示变体的互补优势。大量实验表明,RAMA在基准数据集上表现优异,尤其在基于检索事实证据解决模糊或不合理声明方面表现出色。研究结果强调了整合网络证据与多智能体推理对可信赖多媒体验证的重要性,为更可靠、可扩展的事实核查方案铺平道路。RAMA项目将公开发布于https://github.com/kalendsyang/RAMA.git。
原文摘要 · Abstract (English)
The rapid proliferation of multimodal misinformation presents significant challenges for automated fact-checking systems, especially when claims are ambiguous or lack sufficient context. We introduce RAMA, a novel retrieval-augmented multi-agent framework designed for verifying multimedia misinformation. RAMA incorporates three core innovations: (1) strategic query formulation that transforms multimodal claims into precise web search queries; (2) cross-verification evidence aggregation from diverse, authoritative sources; and (3) a multi-agent ensemble architecture that leverages the complementary strengths of multiple multimodal large language models and prompt variants. Extensive experiments demonstrate that RAMA achieves superior performance on benchmark datasets, particularly excelling in resolving ambiguous or improbable claims by grounding verification in retrieved factual evidence. Our findings underscore the necessity of integrating web-based evidence and multi-agent reasoning for trustworthy multimedia verification, paving the way for more reliable and scalable fact-checking solutions. RAMA will be publicly available at https://github.com/kalendsyang/RAMA.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。