融合视觉与听觉分析,94%准确率检测深度伪造内容
A Multimodal Framework for Deepfake Detection
- 同时分析视频人脸特征与音频频谱图,捕捉多模态异常
- 在真实与伪造音视频混合测试中达到94%识别准确率
- 适合媒体安全、AI伦理及虚假信息防控研究者参考
深度伪造技术的快速发展对数字媒体的真实性构成重大威胁。深度伪造是利用人工智能生成的合成媒体,可逼真地篡改视频和音频,误导现实,带来虚假信息传播、欺诈及个人隐私安全等严重风险。本研究提出一种创新的多模态检测框架,同时针对视觉与听觉特征进行分析。视觉方面,采用先进特征提取技术,捕获九种面部特征,并应用多种机器学习与深度学习模型;音频方面,通过梅尔频谱图提取特征,并使用相应模型进行分析。为实现联合判断,原始数据集中的真实与伪造音视频被互换用于测试,确保样本均衡。基于所提视频与音频分类模型(人工神经网络与VGG19),若任一模态被判定为伪造,则整体样本视为深度伪造。该多模态框架综合视觉与听觉分析,最终达到94%的检测准确率。
原文摘要 · Abstract (English)
The rapid advancement of deepfake technology poses a significant threat to digital media integrity. Deepfakes, synthetic media created using AI, can convincingly alter videos and audio to misrepresent reality. This creates risks of misinformation, fraud, and severe implications for personal privacy and security. Our research addresses the critical issue of deepfakes through an innovative multimodal approach, targeting both visual and auditory elements. This comprehensive strategy recognizes that human perception integrates multiple sensory inputs, particularly visual and auditory information, to form a complete understanding of media content. For visual analysis, a model that employs advanced feature extraction techniques was developed, extracting nine distinct facial characteristics and then applying various machine learning and deep learning models. For auditory analysis, our model leverages mel-spectrogram analysis for feature extraction and then applies various machine learning and deep learningmodels. To achieve a combined analysis, real and deepfake audio in the original dataset were swapped for testing purposes and ensured balanced samples. Using our proposed models for video and audio classification i.e. Artificial Neural Network and VGG19, the overall sample is classified as deepfake if either component is identified as such. Our multimodal framework combines visual and auditory analyses, yielding an accuracy of 94%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。