arXiv:2505.15136cs.SDcs.CR2025-05被引 1

用混合语音数据训练模型,精准识别真人与AI合成语音混合的伪造音频。

Hybrid Audio Detection Using Fine-Tuned Audio Spectrogram Transformers: A Dataset-Driven Evaluation of Mixed AI-Human Speech

  • 基于音谱变换器微调模型,捕捉混合语音中的复杂声学特征。
  • 在混合音频检测任务中达到97%准确率,显著优于现有方法。
  • 适合安全认证、防诈骗系统研发人员参考。

人工智能快速发展催生了先进的音频生成与语音克隆技术,对依赖语音认证的应用构成严重安全威胁。现有数据集和模型多聚焦于区分真人语音与全合成语音,但现实攻击常涉及真人与克隆语音的混合片段。为此,我们构建了一个新型混合音频数据集,包含真人、AI生成、克隆及混合音频样本。进一步提出针对此类复杂声学模式优化的微调音谱变换器(AST)模型。大量实验表明,该方法在混合音频检测任务中显著优于现有基线,分类准确率达97%。研究结果强调了混合数据集与专用模型在提升语音认证系统鲁棒性中的关键作用。

原文摘要 · Abstract (English)

The rapid advancement of artificial intelligence (AI) has enabled sophisticated audio generation and voice cloning technologies, posing significant security risks for applications reliant on voice authentication. While existing datasets and models primarily focus on distinguishing between human and fully synthetic speech, real-world attacks often involve audio that combines both genuine and cloned segments. To address this gap, we construct a novel hybrid audio dataset incorporating human, AI-generated, cloned, and mixed audio samples. We further propose fine-tuned Audio Spectrogram Transformer (AST)-based models tailored for detecting these complex acoustic patterns. Extensive experiments demonstrate that our approach significantly outperforms existing baselines in mixed-audio detection, achieving 97\% classification accuracy. Our findings highlight the importance of hybrid datasets and tailored models in advancing the robustness of speech-based authentication systems.

语音伪造检测混合音频AST模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。