arXiv:2510.09225eess.AScs.CL2025-10中稿 · ICASSP 2026被引 4

语音词段表征差异是无监督分词性能的主要瓶颈。

Unsupervised lexicon learning from speech is limited by representations rather than clustering

  • 用多种自监督语音特征与聚类方法对比实验
  • 发现相同词类内词段表征差异大,制约聚类效果
  • 适合语音处理、无监督学习研究者阅读

零资源词分割与聚类系统旨在无需文本标注的情况下将语音切分为类词单位。尽管已有进展,但生成的词汇表仍不理想。在理想设定下(已知词边界),我们探究性能受限于词段表征还是聚类方法。在英语和汉语数据上,结合多种自监督语音特征(连续/离散、帧级/词级)与不同聚类方法(K-means、层次聚类、图聚类)。最佳系统使用动态时间规整的图聚类搭配连续特征;更快方案采用余弦距离的平均连续特征或离散单元序列的编辑距离。通过控制实验分离表征与聚类因素,结果表明:同词类内词段间的表征变异性——而非聚类方法——是限制性能的主要因素。

原文摘要 · Abstract (English)

Zero-resource word segmentation and clustering systems aim to tokenise speech into word-like units without access to text labels. Despite progress, the induced lexicons are still far from perfect. In an idealised setting with gold word boundaries, we ask whether performance is limited by the representation of word segments, or by the clustering methods that group them into word-like types. We combine a range of self-supervised speech features (continuous/discrete, frame/word-level) with different clustering methods (K-means, hierarchical, graph-based) on English and Mandarin data. The best system uses graph clustering with dynamic time warping on continuous features. Faster alternatives use graph clustering with cosine distance on averaged continuous features or edit distance on discrete unit sequences. Through controlled experiments that isolate either the representations or the clustering method, we demonstrate that representation variability across segments of the same word type -- rather than clustering -- is the primary factor limiting performance.

语音处理无监督学习表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。