用多智能体系统结合大模型,自动验证多媒体真伪。
Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models
- 设计六阶段流程,由专业智能体协同完成验证任务。
- 精准定位内容地理时间信息,追溯跨平台来源。
- 适合媒体审核、舆情监控等真实场景使用。
本文提交至ACMMM25多媒体真实性验证挑战赛。我们构建了一个多智能体验证系统,融合多模态大模型(MLLMs)与专用工具,以识别多媒体虚假信息。系统包含六个阶段:原始数据处理、规划、信息提取、深度调研、证据收集和报告生成。核心深度研究员智能体集成四种工具:反向图像搜索、元数据分析、事实核查数据库及可信新闻处理,用于提取空间、时间、归属与动机上下文。在涉及复杂多媒体内容的挑战数据集样本上,系统成功验证了内容真实性,准确获取地理位置与时间信息,并跨平台追踪来源,有效应对真实世界多媒体验证需求。
原文摘要 · Abstract (English)
This paper presents our submission to the ACMMM25 - Grand Challenge on Multimedia Verification. We developed a multi-agent verification system that combines Multimodal Large Language Models (MLLMs) with specialized verification tools to detect multimedia misinformation. Our system operates through six stages: raw data processing, planning, information extraction, deep research, evidence collection, and report generation. The core Deep Researcher Agent employs four tools: reverse image search, metadata analysis, fact-checking databases, and verified news processing that extracts spatial, temporal, attribution, and motivational context. We demonstrate our approach on a challenge dataset sample involving complex multimedia content. Our system successfully verified content authenticity, extracted precise geolocation and timing information, and traced source attribution across multiple platforms, effectively addressing real-world multimedia verification scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。