用独立成分分析提升语音单元提取质量,显著改善下游识别性能
Discrete Speech Unit Extraction via Independent Component Analysis
- 采用ICA等线性预处理优化自监督语音表示
- 在ASR任务中相较传统方法提升1.8%相对错误率
- 适合关注语音表征压缩与高效建模的研究者
自监督语音模型(S3Ms)已成为语音处理领域的常用工具,其表征可用于下游任务。通过聚类S3M表征可获得离散语音单元(DSUs),作为语音信号的紧凑表示,常用于自动语音识别(ASR)等任务。目前主流方法为k-means聚类,尽管S3M表征具有高维性和冗余性,但针对其预处理以提升聚类质量的研究仍不充分。本文探讨线性预处理方法在提取DSUs中的潜力,评估标准化、主成分分析、白化及独立成分分析(ICA)在基于DSU的ASR基准上的效果,证明其作为k-means预处理的有效性。同时对ICA各分量的正交性与可解释性进行深入分析。
原文摘要 · Abstract (English)
Self-supervised speech models (S3Ms) have become a common tool for the speech processing community, leveraging representations for downstream tasks. Clustering S3M representations yields discrete speech units (DSUs), which serve as compact representations for speech signals. DSUs are typically obtained by k-means clustering. Using DSUs often leads to strong performance in various tasks, including automatic speech recognition (ASR). However, even with the high dimensionality and redundancy of S3M representations, preprocessing S3M representations for better clustering remains unexplored, even though it can affect the quality of DSUs. In this paper, we investigate the potential of linear preprocessing methods for extracting DSUs. We evaluate standardization, principal component analysis, whitening, and independent component analysis (ICA) on DSU-based ASR benchmarks and demonstrate their effectiveness as preprocessing for k-means. We also conduct extensive analyses of their behavior, such as orthogonality or interpretability of individual components of ICA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。