arXiv:2608.24639eess.ASeess.SP2026-08中稿 · IEEE ICASSP 2025被引 11

发现语音中非发声部分对伪造音频检测更有效,准确率提升近50%。

Investigating voiced and unvoiced regions of speech for audio deepfake detection

论文配图:Investigating voiced and unvoiced regions of speech for audio deepfake detection
图 1 · 摘自论文原文
  • 分离语音的发声与非发声区域,分别训练检测模型
  • 非发声区域检测错误率仅6.62%,优于全音频基线
  • 融合两者结果可进一步提升性能,适合音频安全研究者

基于深度神经网络的深度伪造检测系统在基准数据集上已达到高准确率,但多数模型缺乏可解释性,难以向人类评估者提供可信推理。人类常依赖不自然的音高抖动、机械式语调、声学伪影及异常摩擦音等声学线索判断合成音频质量。本研究探讨了语音中发声与非发声区域在区分合成与真实语音中的作用。通过信号周期性度量将语音分解为发声与非发声成分,并分别独立训练基于图注意力的AASIST检测系统。在MLAAD数据集上的实验表明,非发声区域在识别深度伪造语音方面尤为有效,错误率降至6.62%。通过分数级融合发声与非发声区域的结果,整体性能进一步提升,达到5.82%的等错误率,相较使用完整音频的基线系统相对提升49%。

原文摘要 · Abstract (English)

Deep neural network based deepfake detection systems have achieved high levels of accuracy on benchmark datasets and competitions. However, most models lack interpretability. It is challenging to extract reasoning from the network that can convince the human evaluator to trust the decision. Humans often rely on acoustic cues like unnatural pitch jitter, robotic intonation, acoustic artifacts, and unnatural sounding fricatives to judge the quality of the synthetic audio. This study explores the role played by the voiced and unvoiced regions of speech in discriminating synthetic from bonafide speech. A measure of signal periodicity is used to analyze speech into voiced and unvoiced components. Then, the graph attention based AASIST detection system is trained independently on each component. This work compares the accuracy of deepfake detection system using voiced and unvoiced components and analyzes the results on the MLAAD dataset. Our results show that unvoiced regions are particularly more effective in distinguishing synthetic (deepfake) speech from bonafide, and achieves an equal error rate of 6.62%. When combined with voice regions through score-level fusion, the overall performance improves further, yielding a 5.82% EER, a relative improvement of 49% over the baseline system that uses the full audio.

音频伪造检测语音分析深度学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。