arXiv:2512.07351cs.CVcs.AI2025-12被引 3

双代理协同检测深度伪造,提升跨数据集鲁棒性。

DeepAgent: A Dual Stream Multi Agent Fusion for Robust Multimodal Deepfake Detection

  • 分设视觉与音视频不一致检测代理,独立分析多模态信息。
  • 融合模型在跨数据集测试中达97.49%准确率,显著提升鲁棒性。
  • 适合需要高可靠伪造检测的平台与安全系统使用。

合成媒体(尤其是深度伪造)的广泛应用对数字内容验证构成严峻挑战。尽管现有研究结合音频与视觉信息,但多数方法将多模态信息整合于单一模型,仍易受模态错配、噪声和篡改影响。为此,本文提出DeepAgent,一种双流多智能体融合框架,通过并行处理视觉与音频模态实现高效深度伪造检测。该框架包含两个互补智能体:Agent-1采用轻量级基于AlexNet的CNN分析视频,识别深度伪造特征;Agent-2结合声学特征、Whisper生成的语音转录及EasyOCR提取的图像帧文本,检测音视频不一致性。二者决策由随机森林元分类器融合,利用不同决策边界提升最终性能。实验在三个基准数据集上验证,结果表明:Agent-1在合并的Celeb-DF与FakeAVCeleb数据集上测试准确率达94.35%;Agent-2与最终元分类器在FakeAVCeleb上分别达到93.69%和81.56%。跨数据集验证在DeepFakeTIMIT上显示元分类器准确率为97.49%,证明其在多样化数据上的强泛化能力。结果表明,层级融合可有效缓解单模态弱点,验证了多智能体方法在应对多种伪造类型中的有效性。

原文摘要 · Abstract (English)

The increasing use of synthetic media, particularly deepfakes, is an emerging challenge for digital content verification. Although recent studies use both audio and visual information, most integrate these cues within a single model, which remains vulnerable to modality mismatches, noise, and manipulation. To address this gap, we propose DeepAgent, an advanced multi-agent collaboration framework that simultaneously incorporates both visual and audio modalities for the effective detection of deepfakes. DeepAgent consists of two complementary agents. Agent-1 examines each video with a streamlined AlexNet-based CNN to identify the symbols of deepfake manipulation, while Agent-2 detects audio-visual inconsistencies by combining acoustic features, audio transcriptions from Whisper, and frame-reading sequences of images through EasyOCR. Their decisions are fused through a Random Forest meta-classifier that improves final performance by taking advantage of the different decision boundaries learned by each agent. This study evaluates the proposed framework using three benchmark datasets to demonstrate both component-level and fused performance. Agent-1 achieves a test accuracy of 94.35% on the combined Celeb-DF and FakeAVCeleb datasets. On the FakeAVCeleb dataset, Agent-2 and the final meta-classifier attain accuracies of 93.69% and 81.56%, respectively. In addition, cross-dataset validation on DeepFakeTIMIT confirms the robustness of the meta-classifier, which achieves a final accuracy of 97.49%, and indicates a strong capability across diverse datasets. These findings confirm that hierarchy-based fusion enhances robustness by mitigating the weaknesses of individual modalities and demonstrate the effectiveness of a multi-agent approach in addressing diverse types of manipulations in deepfakes.

深度伪造多模态检测智能体协同音视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。