arXiv:2608.12536eess.AS2026-08

用嵌入几何分析语音编码器在印地语中的自发语音检测与合成语音泛化能力

Evaluating Pre-trained Speech Encoders for Spontaneous Speech Detection and Out of Domain Synthetic Speech Generalisation in Indic Languages

论文配图:Evaluating Pre-trained Speech Encoders for Spontaneous Speech Detection and Out of Domain Synthetic Speech Generalisation in Indic Languages
图 1 · 摘自论文原文
  • 对比5个预训练模型在22种印地语上的表现,分析语言区分与语音自然度检测的权衡
  • 发现跨域泛化能力由训练系统与未知合成语音嵌入的接近度决定,而非与真实语音的距离
  • 适合关注深度伪造检测数据选择和模型泛化能力的研究者

基于Transformer的模型在区分自发与剧本语音、真实与合成语音方面表现优异,但这些成果仅限于资源丰富的语言基准,尚未拓展至印地语系列语言,且缺乏对编码器行为的嵌入几何解释或深度伪造泛化失败的分析。本文评估了五个冻结的Transformer编码器(AST、Vaani-FastConformer、Wav2vec2、Whisper、BEATs)在22种印地语上的表现,并开展跨四类TTS系统的多系统泛化实验。除准确率外,引入语言隔离探测与中心点邻近性分析。探测显示编码器在语言区分性和自发性检测间存在依赖模型的权衡;中心点分析表明,跨域泛化能力由训练系统与未见合成语音嵌入的接近度预测,而非其与自然语音的距离,该发现对实际深度伪造检测器的数据选择具有直接指导意义。

原文摘要 · Abstract (English)

Transformer-based models have shown strong accuracy in distinguishing spontaneous from scripted speech and natural from synthetic speech, but these results are established on a narrow set of well-resourced language benchmarks and have not been extended across Indic languages, nor has embedding geometry been used to explain encoder behaviour or deepfake generalisation failure. We address these gaps by evaluating five frozen transformer encoders, AST, Vaani-FastConformer, Wav2vec2, Whisper and BEATs, across 22 Indic languages, and by conducting a multi-system TTS generalisation experiment across four TTS models. Beyond accuracy, we present language isolation probing and centroid proximity analysis. Probing reveals an encoder-dependent trade-off between language-discriminability and spontaneity detection. Centroid analysis shows that out-of-domain generalisation is predicted by a training system's proximity to unseen TTS embeddings, not its distance from natural speech, a finding with direct implications for training data selection in real-world deepfake detectors.

语音识别深度伪造嵌入分析印地语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。