迭代精炼让HuBERT与wav2vec 2.0表现不同,而非训练目标差异。
Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
- 通过多轮伪标签迭代精炼,提升语音表征质量。
- 隐藏表示与词、音素、说话人身份的相关性由迭代次数决定。
- 适合研究自监督语音模型机制的学者参考。
自监督语音表征学习模型因泛用性和下游任务表现优异而广泛应用,但其模型架构对所学语言信息的影响仍研究不足。本研究对比了HuBERT与wav2vec 2.0两个模型,仅最小化改变其架构差异:训练目标与多轮训练中的伪标签迭代精炼。结果表明,隐藏表示与词身份、音素身份、说话人身份的典型相关性差异,由训练迭代次数解释,而非训练目标。研究建议未来探索迭代精炼在编码语言信息中的有效性原因。
原文摘要 · Abstract (English)
Self-supervised models for speech representation learning now see widespread use for their versatility and performance on downstream tasks, but the effect of model architecture on the linguistic information learned in their representations remains under-studied. This study investigates two such models, HuBERT and wav2vec 2.0, and minimally compares two of their architectural differences: training objective and iterative pseudo-label refinement through multiple training iterations. We find that differences in canonical correlation of hidden representations to word identity, phoneme identity, and speaker identity are explained by training iteration, not training objective. We suggest that future work investigate the reason for the effectiveness of iterative refinement in encoding linguistic information in self-supervised speech representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。