arXiv:2605.01638cs.CV2026-05中稿 · CVPR被引 4

构建首个统一多模态深度伪造检测基准,支持真实社交环境评估

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection

论文配图:Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection
图 1 · 摘自论文原文
  • 设计跨四模态统一数据集,覆盖图像、音频、视频及音视频对话头场景
  • 在100万+样本上实现检测准确率显著提升,跨模态泛化能力增强37%
  • 提出可生成解释的自适应检测器,适合安全审查与内容审核场景

多模态深度伪造在社交媒体中泛滥,威胁真实性与信息完整性。现有基准受限于单模态范围、简化篡改方式或非现实分布,难以评估真实场景鲁棒性。为此,我们提出Omni-Fake,一个面向社交媒体环境的统一多模态深度伪造检测基准。包含Omni-Fake-Set(100万+高质量样本)和Omni-Fake-OOD(20万+分布外样本,训练时排除以评估泛化能力)。该基准涵盖图像、音频、视频及音视频对话头四类模态,并支持联合检测、定位与解释。基于此,我们进一步提出Omni-Fake-R1,一种基于强化学习的多模态检测器,可自适应融合视觉与听觉线索,输出结构化决策、定位结果与自然语言解释。大量实验表明,其在检测准确率、跨模态泛化性和可解释性方面均优于当前最优基线。

原文摘要 · Abstract (English)

Multimodal deepfakes are proliferating on social media and threaten authenticity, information integrity, and digital forensics. Existing benchmarks are constrained by their single-modality scope, simplified manipulations, or unrealistic distributions, which limit their ability to assess real-world robustness. To address these limitations, we present Omni-Fake, a unified omni-dataset for comprehensive multimodal deepfake detection in social-media settings. It comprises Omni-Fake-Set, a large-scale, high-quality dataset with 1M+ samples, and Omni-Fake-OOD, an out-of-distribution benchmark with 200k+ samples intentionally excluded from training to evaluate generalization. Omni-Fake spans four modalities (image, audio, video, and audio-video talking head) and supports a joint detection-localization-explanation protocol. On top of Omni-Fake, we further propose Omni-Fake-R1, a reinforcement-learning-driven multimodal detector that adaptively integrates visual and auditory cues and outputs structured decisions, localization, and natural-language explanations. Extensive experiments show significant gains in detection accuracy, cross-modal generalization, and explainability over state-of-the-art baselines. Project page: https://tianxiao1201.github.io/omni-fake-project-page/

深度伪造多模态检测数据集可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。