通过跨模态指纹一致性,让模型不仅能识破假视频,还能定位伪造者。
Attribution-Guided Multimodal Deepfake Detection via Cross-Modal Forensic Fingerprints

- 引入跨模态指纹一致性损失,强制视听特征对齐
- 在FakeAVCeleb上达到99.7%准确率和95.9%伪造源识别率
- 适合需要溯源伪造视频的安防与内容审核场景
音视频深度伪造已达到难以通过感知判断的逼真程度,威胁媒体真实性和生物识别安全。尽管多模态检测有潜力,但多数方法为二分类任务,常依赖数据集特有伪影而非真实的生成痕迹。我们提出归属引导的多模态深度伪造检测框架(AMDD),联合学习检测与伪造源归属。该框架将生成器归属作为结构化正则化,约束共享嵌入空间的几何结构,使其编码生成器特有的取证特征。引入跨模态取证指纹一致性(CMFFC)损失,强制视觉与音频流中由生成器引发的伪影对齐。利用语音与面部动作间的物理耦合关系,合成流程通常破坏此相关性。架构上,采用ResNet50搭配时序注意力处理视觉,配合预训练ResNet18处理梅尔频谱图,弥合了先前模型的编码能力差距。在FakeAVCeleb上,AMDD实现99.7%平衡准确率、99.8% AUC及95.9%归属准确率。跨数据集评估在DeepfakeTIMIT、DFDM和LAV-DF上表明,真实视频检测具有强泛化能力,而对未见生成器的假视频检测仍是开放挑战,本文深入分析其成因。
原文摘要 · Abstract (English)
Audio-visual deepfakes have reached a level of realism that makes perceptual detection unreliable, threatening media integrity and biometric security. While multimodal detection has shown promise, most approaches are binary classification tasks that often latch onto dataset-specific artifacts rather than genuine generative traces. We argue that a detector incapable of identifying how a video was forged is likely learning the wrong signal. Unlike binary detection, attribution-guided learning imposes a stronger geometric constraint on the shared embedding space, forcing the model to encode generator-specific forensic content rather than shortcuts. We propose the Attribution-Guided Multimodal Deepfake Detection (AMDD) framework, which jointly learns to detect and attribute manipulation. AMDD treats generator attribution as a structured regularization that constrains representation geometry toward forensically meaningful features. We introduce a Cross-Modal Forensic Fingerprint Consistency (CMFFC) loss to enforce alignment between generator-induced artifacts in visual and audio streams. This exploits the fact that coherent manipulation leaves correlated traces across modalities, grounded in the physical coupling between speech and facial articulation that synthetic pipelines routinely disrupt. Architecturally, we pair a ResNet50 with temporal attention for visual encoding against a pretrained ResNet18 for mel spectrograms, closing the encoder capacity gap found in prior models. On FakeAVCeleb, AMDD achieves 99.7% balanced accuracy and 99.8% AUC with 95.9% attribution accuracy. Cross-dataset evaluation on DeepfakeTIMIT, DFDM, and LAV-DF confirms that real video detection generalizes robustly, while fake detection on unseen generators remains an open challenge that we analyze in depth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。