arXiv:2510.16489cs.SDeess.AS2025-10被引 1

揭示说话人嵌入空间如何编码声学特征,发现性别影响嵌入机制。

Interpreting the Dimensions of Speaker Embedding Space

  • 用9个可解释声学参数预测嵌入,效果接近7个主成分。
  • 嵌入对男女说话人表现不同,暗示系统隐含性别识别能力。
  • 年龄信息未被有效捕捉,提示嵌入计算可优化。

说话人嵌入广泛应用于说话人验证等任务,将说话人声音特征编码为固定长度向量。这些嵌入常被视为“黑箱”,其与传统声学、语音维度(如年龄、性别)的关系尚未深入研究。本文基于10,000名说话人的大规模语料库,评估三种先进嵌入系统,发现一组9个可解释的声学参数预测嵌入的效果,与7个主成分相当,能解释超过50%的数据方差。结果显示,某些主成分在男女说话人中作用方式不同,表明嵌入系统中存在隐含的性别识别机制。但嵌入未能有效捕捉说话人年龄特征,说明其计算仍有改进空间。

原文摘要 · Abstract (English)

Speaker embeddings are widely used in speaker verification systems and other applications where it is useful to characterise the voice of a speaker with a fixed-length vector. These embeddings tend to be treated as "black box" encodings, and how they relate to conventional acoustic and phonetic dimensions of voices has not been widely studied. In this paper we investigate how state-of-the-art speaker embedding systems represent the acoustic characteristics of speakers as described by conventional acoustic descriptors, age, and gender. Using a large corpus of 10,000 speakers and three embedding systems we show that a small set of 9 acoustic parameters chosen to be "interpretable" predict embeddings about the same as 7 principal components, corresponding to over 50% of variance in the data. We show that some principal dimensions operate differently for male and female speakers, suggesting there is implicit gender recognition within the embedding systems. However we show that speaker age is not well captured by embeddings, suggesting opportunities exist for improvements in their calculation.

说话人识别嵌入分析声学特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。