arXiv:2511.17181cs.CVcs.LG2025-11中稿 · the IEEE/CVF Confe…被引 1

自监督特征可有效检测音视频深度伪造,音频特征表现最佳。

Investigating self-supervised representations for audio-visual deepfake detection

  • 跨模态系统评估自监督特征在音视频中的检测能力
  • 音频特征泛化性最强,达当前最优性能
  • 模型关注语义区域而非虚假噪声,适合真实场景应用

自监督表示在视觉和语音任务中表现优异,但在音视频深度伪造检测中的潜力尚未充分探索。与以往将这些特征单独使用或嵌入复杂架构不同,我们系统评估了它们在音频、视频及多模态下的表现,涵盖唇动和通用视觉内容。从检测效果、编码信息可解释性、跨模态互补性三个维度进行分析。结果表明,多数自监督特征能捕捉深度伪造相关信号,且具有互补性;模型主要关注语义有意义区域,而非伪影(如开头静音)。在所有特征中,音频引导的表示泛化能力最强,达到当前最优结果。但面对真实世界数据仍存在挑战,分析显示该差距源于数据固有难度,而非模型依赖表面模式。项目主页:https://bit-ml.github.io/ssr-dfd。

原文摘要 · Abstract (English)

Self-supervised representations excel at many vision and speech tasks, but their potential for audio-visual deepfake detection remains underexplored. Unlike prior work that uses these features in isolation or buried within complex architectures, we systematically evaluate them across modalities (audio, video, multimodal) and domains (lip movements, generic visual content). We assess three key dimensions: detection effectiveness, interpretability of encoded information, and cross-modal complementarity. We find that most self-supervised features capture deepfake-relevant information, and that this information is complementary. Moreover, models primarily attend to semantically meaningful regions rather than spurious artifacts (such as the leading silence). Among the investigated features, audio-informed representations generalize best and achieve state-of-the-art results. However, generalization to realistic in-the-wild data remains challenging. Our analysis indicates this gap stems from intrinsic dataset difficulty rather than from features latching onto superficial patterns. Project webpage: https://bit-ml.github.io/ssr-dfd.

深度伪造检测自监督学习音视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。