arXiv:2606.10317eess.AScs.SD2026-06中稿 · Interspeech2026

用自监督空间的GMM模型实现可解释的语音转换,提升音色相似度。

SSL-GMMVC: Interpretable Voice Conversion via Locally Linear GMM Transforms in Self-Supervised Representation Space

论文配图:SSL-GMMVC: Interpretable Voice Conversion via Locally Linear GMM Transforms in Self-Supervised Representation Space
图 1 · 摘自论文原文
  • 在自监督表征空间中用高斯混合建模源目标特征,通过后验加权仿射变换实现转换。
  • 混合分量增多时,约束协方差版本超越深度学习基线,音色相似度提升且发音清晰自然。
  • 可解析转换中的旋转与缩放,关联成分选择与音素结构,适合需可解释性的研究者。

我们提出SSL-GMMVC,一种在自监督语音表征空间中的可解释语音转换方法。该方法使用高斯混合模型对源-目标特征对进行建模,并将转换表示为仿射变换的后验加权和。这使得局部线性变换能适应异构的特征空间结构,同时保持解析可计算性。客观与主观评估表明,SSL-GMMVC在保持相当可懂性与自然度的前提下提升了说话人相似度;即使在协方差受限的情况下,随着混合分量数量增加,其性能仍超越深度学习基线。进一步分析揭示了组件选择与音素结构的关联,并展示了学习到的变换中可解释的缩放与旋转行为。这些发现表明,SSL-GMMVC是一种高效且可分析的语音转换框架。

原文摘要 · Abstract (English)

We introduce SSL-GMMVC, an interpretable voice conversion method in self-supervised speech space. The method models paired source-target features with a Gaussian mixture model and performs conversion as a posterior-weighted sum of affine transforms. This yields locally linear transformations that adapt to heterogeneous feature-space structure while remaining analytically tractable. Through objective and subjective evaluations, we show that SSL-GMMVC improves speaker similarity with comparable intelligibility and naturalness, and that even a constrained covariance variant surpasses a deep learning baseline as the number of mixture components increases. Further analyses link component selection to phonetic structure and reveal interpretable scaling and rotation in the learned transforms. These findings highlight SSL-GMMVC as an effective, analyzable framework for voice conversion.

语音转换可解释性自监督GMM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。