arXiv:2503.22503cs.SDcs.CR2025-03被引 1

用Transformer模型检测合成语音,跨技术泛化能力强。

Cross-Technology Generalization in Synthesized Speech Detection: Evaluating AST Models with Modern Voice Generators

  • 基于音频谱图的Transformer架构,融合差异化增强策略
  • 仅用102样本训练,对未知生成器实现3.3% EER
  • 为新型语音合成技术快速适配提供基准,适合安全验证场景

本文评估了音频谱图变换器(AST)架构在合成语音检测中的表现,重点关注其在现代语音生成技术间的泛化能力。通过采用差异化的数据增强策略,模型在对抗ElevenLabs、NotebookLM和Minimax AI等语音生成器时,整体等错误率(EER)达到0.91%。值得注意的是,仅使用单一技术的102个样本进行训练后,模型在完全未见过的语音生成器上仍能达到3.3%的EER。该工作建立了对新兴合成技术快速适应的基准,证明了基于Transformer的架构能够识别不同神经语音合成方法中的共性特征,有助于构建更鲁棒的语音验证系统。

原文摘要 · Abstract (English)

This paper evaluates the Audio Spectrogram Transformer (AST) architecture for synthesized speech detection, with focus on generalization across modern voice generation technologies. Using differentiated augmentation strategies, the model achieves 0.91% EER overall when tested against ElevenLabs, NotebookLM, and Minimax AI voice generators. Notably, after training with only 102 samples from a single technology, the model demonstrates strong cross-technology generalization, achieving 3.3% EER on completely unseen voice generators. This work establishes benchmarks for rapid adaptation to emerging synthesis technologies and provides evidence that transformer-based architectures can identify common artifacts across different neural voice synthesis methods, contributing to more robust speech verification systems.

语音检测Transformer合成语音泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。