分解语音特征,发现不同维度控制音量、音高和发音方式
Interpreting Content and Speaker Characteristics in Factorised Self-Supervised Subspaces
- 用SVD分解WavLM语音特征,分离内容与说话人信息
- 前几维内容空间主控音量、共振峰和发音状态,音高在后几维
- 可精准调节音高和音量,适合语音合成与可控生成研究
自监督语音特征同时包含内容与说话人信息。近期工作采用SVD分解方法,将特征分为共享内容矩阵(捕捉时序变化)和说话人特异性变换(捕捉静态说话人特征)。然而这些成分内部的信息组织方式仍不明确。本文研究了WavLM分解后的内容与说话人子空间维度与语音特征(如音高、音强、发声性)的相关性。结果发现:内容空间的前几维主要编码音强、高阶共振峰和发声性,而音高体现在较后维度;说话人空间最高方差维度强烈关联音高与性别,后续维度捕捉高频变化。干预实验表明,操控这些维度可实现对语音特征的定向控制。联合调整内容与说话人表示,能实现对音高、音强等属性的细粒度调控。
原文摘要 · Abstract (English)
Self-supervised speech features encode both content and speaker information. Recent work introduced an SVD-based factorisation that decomposes these features into a shared content matrix capturing temporal variation and speaker-specific transformations capturing static speaker characteristics. However, how information is organised within these components remains unclear. In this paper, we investigate how the dimensions of WavLM-factorised content and speaker subspaces correlate with speech characteristics such as pitch, intensity, and voicing. We find that leading dimensions in the content space primarily capture intensity, higher-order formants, and voicing, while pitch is encoded in a later dimension. In contrast, the highest-variance speaker dimension is strongly associated with pitch and gender, with later dimensions capturing high-frequency variation. Intervention experiments show that manipulating these dimensions enables targeted control of speech characteristics for speech synthesis. Furthermore, modifying the content and speaker representations jointly provides fine-grained control over characteristics such as pitch and intensity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。