用语音情感特征追踪合成语音来源,效果优于传统方法。
Source Tracing of Synthetic Speech Systems Through Paralinguistic Pre-Trained Representations
- 利用语音情感预训练模型提取合成语音中的独特声调特征。
- 融合两种模型后在合成语音溯源任务上达到新最佳性能。
- 适合关注语音伪造检测与数字身份认证的研究者。
本文聚焦于合成语音生成系统(STSGS)的来源追踪问题。每个生成系统会将其特有的韵律特征——如音高、语调、节奏和语调——嵌入合成语音中,反映其底层生成模型的设计。尽管已有研究使用语音预训练模型(SPTM)的表示,但尚未探索专为韵律处理设计的SPTM在合成语音溯源中的潜力。我们假设:由于能更好捕捉源相关的韵律线索,这类模型的表示更具优势。通过对比多种前沿SPTM(包括韵律、单语、多语及说话人识别模型)的表示,验证了该假设。进一步提出TRIO框架,通过门控机制自适应融合表示,结合典型相关性损失对齐不同表示,并用自注意力进行特征优化。融合TRILLsson(韵律SPTM)与x-vector(说话人识别SPTM)后,TRIO超越单一模型、基线融合方法,成为当前最优方案。
原文摘要 · Abstract (English)
In this work, we focus on source tracing of synthetic speech generation systems (STSGS). Each source embeds distinctive paralinguistic features--such as pitch, tone, rhythm, and intonation--into their synthesized speech, reflecting the underlying design of the generation model. While previous research has explored representations from speech pre-trained models (SPTMs), the use of representations from SPTM pre-trained for paralinguistic speech processing, which excel in paralinguistic tasks like synthetic speech detection, speech emotion recognition has not been investigated for STSGS. We hypothesize that representations from paralinguistic SPTM will be more effective due to its ability to capture source-specific paralinguistic cues attributing to its paralinguistic pre-training. Our comparative study of representations from various SOTA SPTMs, including paralinguistic, monolingual, multilingual, and speaker recognition, validates this hypothesis. Furthermore, we explore fusion of representations and propose TRIO, a novel framework that fuses SPTMs using a gated mechanism for adaptive weighting, followed by canonical correlation loss for inter-representation alignment and self-attention for feature refinement. By fusing TRILLsson (Paralinguistic SPTM) and x-vector (Speaker recognition SPTM), TRIO outperforms individual SPTMs, baseline fusion methods, and sets new SOTA for STSGS in comparison to previous works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。