发现语音模型微调效果受预训练实例影响,不能盲目归因于方法改进。
Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?
- 对比8种微调方法在9个预训练模型上的表现差异。
- 同一方法在不同预训练模型上表现差异显著,最优方案不通用。
- 适合关注微调可靠性与模型选择的语音任务研究者。
监督微调(SFT)被广泛用于将自监督语音表示适配到下游分类任务。单个预训练检查点下观察到的微小提升常被解释为方法层面的改进,即性能上限提高。我们发现这类结论并不可靠,因为SFT结果强烈依赖于具体的预训练实例。我们在3个SUPERB分类任务上系统评估了8种SFT变体,在wav2vec~2.0、HuBERT和WavLM的9个预训练检查点上进行多种子重复实验,使用代表性基础规模模型。结果表明,表现统计上无差别的最优微调方案往往依赖于具体检查点,跨预训练实例的可迁移性有限。这些发现表明,许多报告的下游性能提升反映的是实例与种子相关的激发匹配,而非普遍提升性能上限。
原文摘要 · Abstract (English)
Supervised fine-tuning (SFT) is widely used to adapt self-supervised speech representations to downstream classification tasks. Small gains observed under a single pretrained checkpoint are often interpreted as method-level improvements, i.e., a higher attainable performance ceiling. We show that such conclusions are not always reliable because SFT outcomes depend strongly on the specific pretrained instance. We conduct a systematic study on 3 SUPERB classification tasks, evaluating 8 SFT variants across 9 pretrained checkpoints from wav2vec~2.0, HuBERT, and WavLM, with multi-seed repetitions on representative base-scale models. We find that the identity of the statistically indistinguishable top-group SFT recipe is often checkpoint-dependent, with limited transferability across pretrained instances. These findings suggest that many reported downstream gains reflect instance and seed dependent elicitation match, rather than universally improving the attainable performance ceiling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。