用多个专家智能体协作检测深度伪造视频,效果超越闭源大模型。
Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

- 设计四类专家智能体从纹理、光照、运动、物理四方面分析伪造痕迹。
- 在10万视频数据集上,小模型组合比闭源大模型更准确,跨域检测领先。
- 适用于需要高可解释性与强泛化能力的AI安全场景,如内容审核。
生成式人工智能制造高度逼真的深度伪造视频引发严重伦理问题并威胁AI安全。现有深度伪造视频基准覆盖近期合成方法有限,且缺乏可靠细粒度文本标注。传统检测器和多模态大语言模型(MLLM)通常依赖单一模型或视角,难以捕捉细微伪造痕迹,泛化能力受限。为此,我们构建了包含10万视频、涵盖33种合成方法(包括Seedance 2.0)的大型数据集FaceVid-Forensics-100K,提供细粒度视觉观察与一致结论的法证解释标注,通过先进MLLM驱动的多模型聚合与冲突解决管道自动生成。基于此基准,我们提出多智能体法证推理框架:四个领域专家智能体分别从纹理、光照、运动、物理角度独立分析伪造线索,一名裁判智能体整合报告生成最终预测与解释。在跨域测试集上的大量评估显示,尽管全部由小型开源MLLM构成,该框架仍优于所有对比方法,包括闭源GPT和Gemini模型,并在各项指标上排名第一。
原文摘要 · Abstract (English)
The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmarks provide limited coverage of recent synthesis methods and generally lack reliable fine-grained textual annotations. Meanwhile, conventional detectors and multimodal large language models (MLLMs), whether operating as a single model or relying on a single analytical perspective, often fail to capture subtle forgery artifacts, limiting their generalization to emerging AI-generated methods. To address these limitations, we introduce FaceVid-Forensics-100K, a large-scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire-face synthesis, including recent generators such as Seedance 2.0. The dataset provides fine-grained textual annotations of visual observations and verdict-consistent forensic explanations, automatically synthesized through a multi-model aggregation and conflict-resolution pipeline powered by advanced MLLMs. Building on this benchmark, we propose a multi-agent forensic reasoning framework that employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives: texture, lighting, motion, and physics. A judge agent then reconciles their reports to produce a final prediction together with an explanation. Extensive evaluations on out-of-domain test sets show that, despite being composed entirely of small open-source MLLMs, our framework outperforms all methods including closed-source GPT and Gemini models and ranks first across all reported metrics on this benchmark. The project page is available at https://xavierjiezou.github.io/ARGUS/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。