arXiv:2603.03096eess.AScs.CL2026-03

发现自监督语音模型中,各维度可独立控制音高、音量等说话人特征。

Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features

  • 用主成分分析揭示语音特征在自监督模型中的维度分布。
  • 主成分维度分别对应音高、音量、噪声、共振峰等特征。
  • 通过操控特定维度可独立改变说话人特征,适用于语音编辑应用。

自监督语音模型如何组织其表征?以往研究关注信息在不同层的编码方式,但较少探讨语音特征是否存在于单个特征维度中。本文通过对语音片段平均表示进行主成分分析,专门考察说话人特征。在多种自监督语音模型中,我们发现解释最大方差的主成分编码了音高及性别等关联特征;其他主成分则与强度、噪声水平、第二共振峰及高频特征相关。合成分析显示,这些特征对应的维度相互独立,且通过操控相应维度可实现特征的可控修改。

原文摘要 · Abstract (English)

How do speech models trained through self-supervised learning structure their representations? Previous studies have looked at how information is encoded in feature vectors across different layers. But few studies have considered whether speech characteristics are captured within individual dimensions of SSL features. In this paper we specifically look at speaker information using PCA on utterance-averaged representations. For a range of SSL models, we find that the principal dimension that explains most variance encodes pitch and associated characteristics like gender. Other individual principal dimensions correlate with intensity, noise levels, the second formant, and higher frequency characteristics. We then use synthesis analyses to show that the dimensions for most characteristics are isolated from each other's influence. We further show that characteristics can be changed by manipulating the corresponding dimensions.

语音表征自监督学习特征解耦语音编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。