用简单声学特征检测语音克隆失败,快速识别劣质合成语音
Low-Cost Detection of Degraded Voice Clones via Source-Output Acoustic Consistency

- 基于源-滤波模型,用基频和噪声比等声学特征做一致性检测
- WaveRNN下准确率达85.2%,HiFi-GAN下达80.0%,效果优于声道长度
- 特征互补,适合临床等需快速筛查的高可靠性场景
生成式语音技术的进步使得自动检测明显失败的合成输出变得尤为重要,尤其在AVATAR疗法等临床场景中,精神分裂症患者与幻觉声音的计算机生成代表互动,劣质合成可能破坏沉浸感与治疗参与度。本文研究低维、可解释的源-输出声学特征能否作为轻量级首检工具。受语音源-滤波模型启发,测试了中位基频(f0)作为源相关一致性度量,对比声道长度(VTL)作为滤波相关度量,以及谐波-噪声比(HNR)作为噪声相关描述符。使用两种声码器家族(WaveRNN, n=54;HiFi-GAN, n=40)生成的人工标注语音克隆样本,通过输入-输出特征空间的非对称阈值法评估。WaveRNN下f0与HNR均达85.2%准确率,优于VTL的64.8%;HiFi-GAN下HNR为80.0%,f0为77.5%,VTL为67.5%。样本级重叠与谱图检查显示,f0与HNR捕捉部分不同故障模式,而非冗余排序。结果表明,简单的源-输出声学一致性测量可有效实现劣质语音克隆的首检,支持在需快速拒绝失败合成语音的应用中采用可解释的阈值筛选。
原文摘要 · Abstract (English)
Recent advances in generative speech have increased the need for automatic detection of obviously failed synthetic outputs. This is particularly important in clinical settings such as AVATAR therapy, in which schizophrenia patients engage with a computer-generated representation of their hallucinated voices and degraded synthesis may disrupt immersion and therapeutic engagement. We investigate whether low-dimensional, interpretable source-output acoustic features can provide a lightweight first-pass detector of degraded voice-cloning outputs. Motivated by source-filter models of speech, we first test median fundamental frequency (f0) as a source-related consistency measure, and compare it with vocal tract length (VTL) as a filter-related measure and Harmonics-to-Noise Ratio (HNR) as a noise-related descriptor. Human-labeled voice-cloning samples generated with two vocoder families, WaveRNN (n=54) and HiFi-GAN (n=40), were evaluated using an asymmetric thresholding procedure in the input-output feature space. For WaveRNN, f0 and HNR both achieved 85.2% accuracy, outperforming VTL (64.8%). For HiFi-GAN, HNR achieved 80.0% accuracy, followed by f0 at 77.5% and VTL at 67.5%. Sample-level overlap and spectrographic inspection showed that f0 and HNR capture partly distinct failure patterns, rather than providing redundant rankings of the same samples. These results show that simple source-output acoustic consistency measures can provide useful first-pass detection of degraded voice clones, and support the use of interpretable threshold-based screening in applications where failed synthetic speech must be rejected quickly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。