用随机采样优化音视频伪造检测模型,提升准确率与泛化能力
Gumbel Rao Monte Carlo based Bi-Modal Neural Architecture Search for Audio-Visual Deepfake Detection
- 基于戈布尔-罗亚蒙特卡洛采样,改进融合机制稳定性
- 在FakeAVCeleb上达95.4% AUC,参数量极低
- 适合需要轻量化高精度检测的场景
深度伪造对生物特征认证系统构成严重威胁,生成高度逼真的合成媒体。现有多模态深度伪造检测方法难以适应多样数据,且依赖简单融合方式。为此,我们提出一种基于戈布尔-罗亚蒙特卡洛采样的双模态神经架构搜索框架(GRMC-BMNAS),通过戈布尔-罗亚蒙特卡洛采样优化多模态融合。该方法通过罗亚-布莱克韦尔化降低直通戈布尔软最大值(STGS)的方差,提升训练稳定性。采用两级搜索策略,优化网络结构、参数与性能。关键特征从主干网络中高效提取,细胞结构内采用加权融合操作整合多源信息。通过调节温度和蒙特卡洛样本数,获得最大化分类性能与更好泛化能力的架构。在FakeAVCeleb和SWAN-DF数据集上的实验表明,该方法以极少模型参数实现95.4%的AUC。
原文摘要 · Abstract (English)
Deepfakes pose a critical threat to biometric authentication systems by generating highly realistic synthetic media. Existing multimodal deepfake detectors often struggle to adapt to diverse data and rely on simple fusion methods. To address these challenges, we propose Gumbel-Rao Monte Carlo Bi-modal Neural Architecture Search (GRMC-BMNAS), a novel architecture search framework that employs Gumbel-Rao Monte Carlo sampling to optimize multimodal fusion. It refines the Straight through Gumbel Softmax (STGS) method by reducing variance with Rao-Blackwellization, stabilizing network training. Using a two-level search approach, the framework optimizes the network architecture, parameters, and performance. Crucial features are efficiently identified from backbone networks, while within the cell structure, a weighted fusion operation integrates information from various sources. By varying parameters such as temperature and number of Monte carlo samples yields an architecture that maximizes classification performance and better generalisation capability. Experimental results on the FakeAVCeleb and SWAN-DF datasets demonstrate an impressive AUC percentage of 95.4\%, achieved with minimal model parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。