arXiv:2509.00230cs.SDcs.AI2025-09被引 1

对比三类语音模型在说话人识别中的表现,发现早期层更有效。

Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks

  • 用层间分析方法研究模型内部表示
  • Wav2Vec 2.0和XLS-R早层表现更优,Whisper深层更强
  • 确定各模型微调时最优层数,适合语音识别研究者

本研究评估了三种先进语音编码模型——Wav2Vec 2.0、XLS-R 和 Whisper——在说话人识别任务中的表现。通过微调这些模型,并利用 SVCCA、k-means 聚类和 t-SNE 可视化分析其逐层表征,发现 Wav2Vec 2.0 和 XLS-R 在早期层能有效捕捉说话人特定特征,微调后性能与稳定性均提升;Whisper 在深层表现更佳。此外,本文还确定了各模型在微调用于说话人识别任务时的最优 Transformer 层数量。

原文摘要 · Abstract (English)

This study evaluates the performance of three advanced speech encoder models, Wav2Vec 2.0, XLS-R, and Whisper, in speaker identification tasks. By fine-tuning these models and analyzing their layer-wise representations using SVCCA, k-means clustering, and t-SNE visualizations, we found that Wav2Vec 2.0 and XLS-R capture speaker-specific features effectively in their early layers, with fine-tuning improving stability and performance. Whisper showed better performance in deeper layers. Additionally, we determined the optimal number of transformer layers for each model when fine-tuned for speaker identification tasks.

说话人识别语音编码Transformer模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。