自监督语音模型通过正交子空间编码前后音素上下文。
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
- 单帧表示中叠加前、当前、后音素的音位向量。
- 不同位置音素向量间保持正交性,形成隐式音素边界。
- 揭示了语音模型如何以组合方式编码上下文信息。
基于Transformer的自监督语音模型(S3Ms)常被称为上下文相关,但其具体含义尚不明确。本文研究单帧级S3M表示如何编码音素及其周围上下文。已有研究表明,S3M以组合方式表示音素,例如[ m ]的表示中叠加了清浊、双唇、鼻音等音位向量。本文进一步提出,相邻音素序列的音位信息也以组合方式编码于单帧中,即前、当前、后音素对应的向量在单帧表示内叠加。我们发现该结构具有相对位置间的正交性,以及隐式音素边界的出现。这些发现深化了对上下文依赖型S3M表示机制的理解。
原文摘要 · Abstract (English)
Transformer-based self-supervised speech models (S3Ms) are often described as contextualized, yet what this entails remains unclear. Here, we focus on how a single frame-level S3M representation can encode phones and their surrounding context. Prior work has shown that S3Ms represent phones compositionally; for example, phonological vectors such as voicing, bilabiality, and nasality vectors are superposed in the S3M representation of [m]. We extend this view by proposing that phonological information from a sequence of neighboring phones is also compositionally encoded in a single frame, such that vectors corresponding to previous, current, and next phones are superposed within a single frame-level representation. We show that this structure has several properties, including orthogonality between relative positions, and emergence of implicit phonetic boundaries. Together, our findings advance our understanding of context-dependent S3M representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。