研究语音表示的各向异性对关键词识别的影响,发现预训练模型仍能有效识别无转录词。
Assessing the Impact of Anisotropy in Neural Representations of Speech: A Case Study on Keyword Spotting
- 用动态时间规整评估语音嵌入的相似性
- 即使存在各向异性,模型仍可准确识别关键词
- 适合关注语音表示鲁棒性的研究者
预训练语音表示如wav2vec2和HuBERT表现出显著的各向异性,导致随机嵌入间相似性过高。尽管这一现象广泛存在,其对下游任务的影响尚不明确。本文针对计算文献语言学中的关键词识别任务,利用动态时间规整(Dynamic Time Warping)评估了各向异性的影响。结果表明,尽管存在各向异性,wav2vec2的相似性度量仍能有效识别未标注词汇。该发现凸显了这些表示在捕捉语音结构和跨说话人泛化方面的鲁棒性,强调了预训练在学习丰富且不变语音表征中的关键作用。
原文摘要 · Abstract (English)
Pretrained speech representations like wav2vec2 and HuBERT exhibit strong anisotropy, leading to high similarity between random embeddings. While widely observed, the impact of this property on downstream tasks remains unclear. This work evaluates anisotropy in keyword spotting for computational documentary linguistics. Using Dynamic Time Warping, we show that despite anisotropy, wav2vec2 similarity measures effectively identify words without transcription. Our results highlight the robustness of these representations, which capture phonetic structures and generalize across speakers. Our results underscore the importance of pretraining in learning rich and invariant speech representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。