自监督语音模型能用可解释的向量算术表示语音音位特征。
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
- 在96种语言中发现语音表征空间存在对应音位特征的线性方向。
- 音位向量的长度与对应音位在语音中的实现程度呈连续相关。
- 通过向量运算可生成语音变体,适合语音处理与认知研究者。
自监督语音模型(S3Ms)虽已知编码丰富的语音信息,但其内部结构仍不明确。本文在96种语言上开展全面研究,分析S3M表征的底层结构,重点关注音位向量。结果表明,模型表征空间中存在对应音位特征的线性方向;这些音位向量的尺度与对应音位特征的声学实现程度呈连续相关。例如,[d]与[t]之差生成一个清浊音向量:将该向量加至[p]可得到[b],缩放则产生清浊度连续变化的语音。这些发现表明,S3Ms以可解释且可组合的向量形式编码语音,展现出音位向量算术特性。所有代码与交互演示已公开于https://github.com/juice500ml/phonetic-arithmetic。
原文摘要 · Abstract (English)
Self-supervised speech models (S3Ms) are known to encode rich phonetic information, yet how this information is structured remains underexplored. We conduct a comprehensive study across 96 languages to analyze the underlying structure of S3M representations, with particular attention to phonological vectors. We first show that there exist linear directions within the model's representation space that correspond to phonological features. We further demonstrate that the scale of these phonological vectors correlate to the degree of acoustic realization of their corresponding phonological features in a continuous manner. For example, the difference between [d] and [t] yields a voicing vector: adding this vector to [p] produces [b], while scaling it results in a continuum of voicing. Together, these findings indicate that S3Ms encode speech using phonologically interpretable and compositional vectors, demonstrating phonological vector arithmetic. All code and interactive demos are available at https://github.com/juice500ml/phonetic-arithmetic .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。