对比三类语音模型在说话人识别中的表现,发现早期层更有效。
Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
- 用层间分析方法研究模型内部表示
- Wav2Vec 2.0和XLS-R早层表现更优,Whisper深层更强
- 确定各模型微调时最优层数,适合语音识别研究者
本研究评估了三种先进语音编码模型——Wav2Vec 2.0、XLS-R 和 Whisper——在说话人识别任务中的表现。通过微调这些模型,并利用 SVCCA、k-means 聚类和 t-SNE 可视化分析其逐层表征,发现 Wav2Vec 2.0 和 XLS-R 在早期层能有效捕捉说话人特定特征,微调后性能与稳定性均提升;Whisper 在深层表现更佳。此外,本文还确定了各模型在微调用于说话人识别任务时的最优 Transformer 层数量。
原文摘要 · Abstract (English)
This study evaluates the performance of three advanced speech encoder models, Wav2Vec 2.0, XLS-R, and Whisper, in speaker identification tasks. By fine-tuning these models and analyzing their layer-wise representations using SVCCA, k-means clustering, and t-SNE visualizations, we found that Wav2Vec 2.0 and XLS-R capture speaker-specific features effectively in their early layers, with fine-tuning improving stability and performance. Whisper showed better performance in deeper layers. Additionally, we determined the optimal number of transformer layers for each model when fine-tuned for speaker identification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。