系统梳理跨模态深度伪造检测方法与挑战
Passive Deepfake Detection Across Multi-modalities: A Comprehensive Survey
- 从图像、视频、音频等多模态角度分析被动检测思路
- 评估模型在新生成技术下的泛化能力与抗干扰性能
- 适合安全研究者与AI伦理领域从业者参考
近年来,深度伪造(DF)被用于个人冒充、虚假信息传播及艺术家风格模仿等恶意目的,引发伦理与安全担忧。本文全面综述并对比了跨多模态(图像、视频、音频及多模态)的被动深度伪造检测方法,探讨模态间的相互关系。除检测准确率外,还扩展分析了真实部署中关键性能维度:对新型生成技术的泛化能力、对抗性篡改与后期处理的鲁棒性、生成源定位精度,以及在真实运行环境下的稳定性。此外,分析了现有数据集、基准测试与评估指标的优劣。最后提出未来研究方向,以应对该领域尚未探索和新兴的问题。本综述为研究人员与实践者提供了理解当前进展、方法体系及未来趋势的综合性资源。
原文摘要 · Abstract (English)
In recent years, deepfakes (DFs) have been utilized for malicious purposes, such as individual impersonation, misinformation spreading, and artists style imitation, raising questions about ethical and security concerns. In this survey, we provide a comprehensive review and comparison of passive DF detection across multiple modalities, including image, video, audio, and multi-modal, to explore the inter-modality relationships between them. Beyond detection accuracy, we extend our analysis to encompass crucial performance dimensions essential for real-world deployment: generalization capabilities across novel generation techniques, robustness against adversarial manipulations and postprocessing techniques, attribution precision in identifying generation sources, and resilience under real-world operational conditions. Additionally, we analyze the advantages and limitations of existing datasets, benchmarks, and evaluation metrics for passive DF detection. Finally, we propose future research directions that address these unexplored and emerging issues in the field of passive DF detection. This survey offers researchers and practitioners a comprehensive resource for understanding the current landscape, methodological approaches, and promising future directions in this rapidly evolving field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。