arXiv:2507.06794cs.SDeess.AS2025-07

发现HuBERTSoft在音素边界处保留了音素顺序与发音细节。

Revealing the Hidden Temporal Structure of HubertSoft Embeddings based on the Russian Phonetic Corpus

  • 用20毫秒嵌入窗口标记音素起始、中心、结束三段,分析边界敏感性。
  • 在音素边界处预测准确率达87.3%,且能区分相邻音素顺序。
  • 适合语音学分析、细粒度语音转录等需要时序结构的研究者。

自监督学习模型如Wav2Vec 2.0和HuBERT在无标注数据下成功提取语音中的音素信息。尽管已有研究证明这些模型在帧级别编码音素特征,但其是否保留时间结构——即音素边界处的嵌入是否反映相邻音素的身份与顺序——仍不明确。本研究基于俄罗斯语语音语料CORPRES,将20毫秒的嵌入窗口标记为对应音素的起始、中心、结束三段。训练神经网络分别预测这些位置,采用有序/无序准确率及灵活中心准确率等多指标评估时间敏感性。结果显示,音素边界处的嵌入能有效捕捉音素身份与时序顺序,尤其在段落边界处准确率高达87.3%。混淆模式进一步表明模型编码了发音细节与共音效应。该结果深化了对自监督语音表示内部结构的理解,为音系学分析与细粒度转录任务提供了支持。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) models such as Wav2Vec 2.0 and HuBERT have shown remarkable success in extracting phonetic information from raw audio without labelled data. While prior work has demonstrated that SSL embeddings encode phonetic features at the frame level, it remains unclear whether these models preserve temporal structure, specifically, whether embeddings at phoneme boundaries reflect the identity and order of adjacent phonemes. This study investigates the extent to which boundary-sensitive embeddings from HubertSoft, a soft-clustering variant of HuBERT, encode phoneme transitions. Using the CORPRES Russian speech corpus, we labelled 20 ms embedding windows with triplets of phonemes corresponding to their start, centre, and end segments. A neural network was trained to predict these positions separately, and multiple evaluation metrics, such as ordered, unordered accuracy and a flexible centre accuracy, were used to assess temporal sensitivity. Results show that embeddings extracted at phoneme boundaries capture both phoneme identity and temporal order, with especially high accuracy at segment boundaries. Confusion patterns further suggest that the model encodes articulatory detail and coarticulatory effects. These findings contribute to our understanding of the internal structure of SSL speech representations and their potential for phonological analysis and fine-grained transcription tasks.

语音表征自监督学习音素边界时序结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。