提出混合语音伪造检测新数据集,揭示微调无法应对复杂攻击。
When Fine-Tuning is Not Enough: Lessons from HSAD on Hybrid and Adversarial Audio Spoof Detection
- 构建含4类语音的HSAD数据集,涵盖真实、克隆、生成及混合语音
- 预训练模型在混合攻击下性能崩溃,微调后仍难识别未见组合
- 专用适配可提升准确率至97%以上,适合安全系统开发者参考
人工智能快速发展使语音合成与克隆高度逼真,威胁语音认证、智能助手和电信安全。现有研究多将伪造检测视为二分类任务,但现实攻击常涉及真实与合成语音混合,检测难度显著增加。为此,我们提出混合伪造音频数据集HSAD,包含1,248条纯净语句和41,044条退化语句,分为四类:人声、克隆、零样本生成及混合语音。每条样本标注伪造方法、说话人身份与退化元信息,支持细粒度分析。我们评估六种基于Transformer的模型,包括频谱编码器(MIT-AST、MattyB95-AST)和自监督波形模型(Wav2Vec2、HuBERT)。结果表明:预训练模型在混合条件下过度泛化并失效;针对伪造场景的微调虽提升可分性,但仍难以处理未见组合;在HSAD上进行数据集特定适配后,模型性能显著提升(AST准确率超97%,F1约99%),但复杂混合仍存残余错误。这些发现表明,仅靠微调不足——如HSAD这类具备对抗意识的基准数据集对暴露模型校准失败、偏见及影响因素至关重要。因此,HSAD不仅提供数据,更构建了增强语音认证系统鲁棒性的分析框架。
原文摘要 · Abstract (English)
The rapid advancement of AI has enabled highly realistic speech synthesis and voice cloning, posing serious risks to voice authentication, smart assistants, and telecom security. While most prior work frames spoof detection as a binary task, real-world attacks often involve hybrid utterances that mix genuine and synthetic speech, making detection substantially more challenging. To address this gap, we introduce the Hybrid Spoofed Audio Dataset (HSAD), a benchmark containing 1,248 clean and 41,044 degraded utterances across four classes: human, cloned, zero-shot AI-generated, and hybrid audio. Each sample is annotated with spoofing method, speaker identity, and degradation metadata to enable fine-grained analysis. We evaluate six transformer-based models, including spectrogram encoders (MIT-AST, MattyB95-AST) and self-supervised waveform models (Wav2Vec2, HuBERT). Results reveal critical lessons: pretrained models overgeneralize and collapse under hybrid conditions; spoof-specific fine-tuning improves separability but struggles with unseen compositions; and dataset-specific adaptation on HSAD yields large performance gains (AST greater than 97 percent and F1 score is approximately 99 percent), though residual errors persist for complex hybrids. These findings demonstrate that fine-tuning alone is not sufficient-robust hybrid-aware benchmarks like HSAD are essential to expose calibration failures, model biases, and factors affecting spoof detection in adversarial environments. HSAD thus provides both a dataset and an analytic framework for building resilient and trustworthy voice authentication systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。