自动调整阈值,让伪造语音模型识别更准更可靠。
Reject Threshold Adaptation for Open-Set Model Attribution of Deepfake Audio
- 用重建误差构建匹配与不匹配样本分布
- 通过概率最小化计算自适应拒绝阈值
- 解决不同数据分布下阈值失效问题,适合实际应用
面向开放环境的深度伪造语音模型归属识别是新兴研究方向,旨在识别伪造语音的生成模型。以往方法需手动设定未知类别的拒绝阈值以比较预测概率,但模型常过拟合训练样本,产生过度自信的预测;且当前数据集有效的阈值在其他数据分布下可能失效。为此,本文提出一种具备拒绝阈值自适应能力的框架(ReTA)。具体地,重建误差学习模块通过结合系统指纹表示与目标类别或随机选取的其他类别标签进行训练,生成匹配与不匹配的重构样本,建立各分类的重建误差分布,为拒绝阈值计算模块奠定基础。该模块采用高斯概率估计拟合匹配与不匹配重建误差分布,并通过概率最小化准则计算所有类别的自适应拒绝阈值。实验表明,ReTA能有效提升深度伪造语音的开放集模型归属识别性能。
原文摘要 · Abstract (English)
Open environment oriented open set model attribution of deepfake audio is an emerging research topic, aiming to identify the generation models of deepfake audio. Most previous work requires manually setting a rejection threshold for unknown classes to compare with predicted probabilities. However, models often overfit training instances and generate overly confident predictions. Moreover, thresholds that effectively distinguish unknown categories in the current dataset may not be suitable for identifying known and unknown categories in another data distribution. To address the issues, we propose a novel framework for open set model attribution of deepfake audio with rejection threshold adaptation (ReTA). Specifically, the reconstruction error learning module trains by combining the representation of system fingerprints with labels corresponding to either the target class or a randomly chosen other class label. This process generates matching and non-matching reconstructed samples, establishing the reconstruction error distributions for each class and laying the foundation for the reject threshold calculation module. The reject threshold calculation module utilizes gaussian probability estimation to fit the distributions of matching and non-matching reconstruction errors. It then computes adaptive reject thresholds for all classes through probability minimization criteria. The experimental results demonstrate the effectiveness of ReTA in improving the open set model attributes of deepfake audio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。