arXiv:2510.06706cs.SDcs.CL2025-10中稿 · 2025 IEEE Internat…被引 2

用新型神经网络提升语音伪造检测准确率

XLSR-Kanformer: A KAN-Intergrated model for Synthetic Speech Detection

  • 用KAN替代传统MLP,增强模型表达能力
  • 在ASVspoof2021上EER降低60.55%,达0.70%
  • 适配多种自监督模型,通用性强

近期语音合成技术的进步带来了日益复杂的伪造攻击,给自动说话人验证系统带来挑战。尽管基于自监督学习(SSL)的XLSR-Conformer模型在语音伪造检测中表现优异,但仍有改进空间。本文提出将XLSR-Conformer中的多层感知机(MLP)替换为基于柯尔莫哥洛夫-阿诺德表示定理的柯尔莫哥洛夫-阿诺德网络(KAN),这是一种强大的通用逼近器。在ASVspoof2021数据集上的实验表明,该改进使模型在LA和DF测试集上的等错误率(EER)相对提升60.55%,并在21LA集上达到0.70%的EER。此外,该替换对多种SSL架构均具鲁棒性。结果表明,将KAN融入基于SSL的模型是提升语音伪造检测性能的可行方向。

原文摘要 · Abstract (English)

Recent advancements in speech synthesis technologies have led to increasingly sophisticated spoofing attacks, posing significant challenges for automatic speaker verification systems. While systems based on self-supervised learning (SSL) models, particularly the XLSR-Conformer architecture, have demonstrated remarkable performance in synthetic speech detection, there remains room for architectural improvements. In this paper, we propose a novel approach that replaces the traditional Multi-Layer Perceptron (MLP) in the XLSR-Conformer model with a Kolmogorov-Arnold Network (KAN), a powerful universal approximator based on the Kolmogorov-Arnold representation theorem. Our experimental results on ASVspoof2021 demonstrate that the integration of KAN to XLSR-Conformer model can improve the performance by 60.55% relatively in Equal Error Rate (EER) LA and DF sets, further achieving 0.70% EER on the 21LA set. Besides, the proposed replacement is also robust to various SSL architectures. These findings suggest that incorporating KAN into SSL-based models is a promising direction for advances in synthetic speech detection.

语音伪造KANSSL检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。