提出解耦特征提取框架,提升语音检测模型跨域泛化能力。
Improving Generalization for AI-Synthesized Voice Detection
- 通过解耦技术分离声码器相关伪影特征,实现域无关表示
- 在跨域检测中误差率降低7.59%,内域检测提升5.12%
- 适合需要长期稳定检测合成语音的应用场景
AI合成语音技术虽具广泛应用潜力,但也存在被滥用风险。现有检测模型在同域评估中表现良好,但在跨域泛化方面仍面临挑战,难以应对新型语音生成器。当前方法依赖预设声码器,且易受背景噪声和说话人身份影响。本文提出一种创新的解耦框架,用于提取与声码器相关的域无关伪影特征。利用这些特征,模型可在平坦损失曲面中学习,避免陷入次优解,从而提升泛化性能。大量实验表明,该方法在基准测试中优于现有最优模型:内域评估中等错误率(EER)降低5.12%,跨域评估中提升7.59%。
原文摘要 · Abstract (English)
AI-synthesized voice technology has the potential to create realistic human voices for beneficial applications, but it can also be misused for malicious purposes. While existing AI-synthesized voice detection models excel in intra-domain evaluation, they face challenges in generalizing across different domains, potentially becoming obsolete as new voice generators emerge. Current solutions use diverse data and advanced machine learning techniques (e.g., domain-invariant representation, self-supervised learning), but are limited by predefined vocoders and sensitivity to factors like background noise and speaker identity. In this work, we introduce an innovative disentanglement framework aimed at extracting domain-agnostic artifact features related to vocoders. Utilizing these features, we enhance model learning in a flat loss landscape, enabling escape from suboptimal solutions and improving generalization. Extensive experiments on benchmarks show our approach outperforms state-of-the-art methods, achieving up to 5.12% improvement in the equal error rate metric in intra-domain and 7.59% in cross-domain evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。