用新型神经网络提升语音伪造检测准确率
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
- 用柯尔莫哥洛夫-阿诺德网络替代传统模型中的全连接层
- 在ASVspoof2021数据集上相对性能提升60.55%,最低错误率0.70%
- 适合关注语音安全与深度伪造防御的研究者
语音合成技术的进步带来了更复杂的欺骗攻击,对自动说话人验证系统构成严峻挑战。尽管基于自监督学习(SSL)的XLSR-Conformer模型已在语音伪造检测中表现出色,但其架构仍有优化空间。本文提出将传统多层感知机替换为基于柯尔莫哥洛夫-阿诺德表示定理的新型网络结构KAN。在ASVspoof2021数据集上的实验表明,将KAN融入SSL模型可使LA和DF集的性能相对提升60.55%,并在21LA集上达到0.70%的等错误率(EER)。结果表明,将KAN引入基于SSL的模型是提升语音伪造检测能力的重要方向。
原文摘要 · Abstract (English)
Recent advancements in speech synthesis technologies have led to increasingly advanced spoofing attacks, posing significant challenges for automatic speaker verification systems. While systems based on self-supervised learning (SSL) models, particularly the XLSR-Conformer model, have demonstrated remarkable performance in synthetic speech detection, there remains room for architectural improvements. In this paper, we propose a novel approach that replaces the traditional Multi-Layer Perceptron in the XLSR-Conformer model with a Kolmogorov-Arnold Network (KAN), a novel architecture based on the Kolmogorov-Arnold representation theorem. Our results on ASVspoof2021 demonstrate that integrating KAN into the SSL-based models can improve the performance by 60.55% relatively on LA and DF sets, further achieving 0.70% EER on the 21LA set. These findings suggest that incorporating KAN into SSL-based models is a promising direction for advances in synthetic speech detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。