用德语字幕微调Whisper,发现基准测试被污染,真实错误率仅为测量值的三分之一。
Subtitle-Aligned Fine-Tuning of Whisper for Swiss German ASR: Benchmark Contamination, Convention Mismatch, and an Honest Baseline at 25.6% WER (13.8% cWER)
- 用广播音频和德语字幕弱监督微调Whisper大模型
- 最佳模型在瑞士德语测试集上达25.6% WER(cWER 13.8%)
- 揭示基准污染问题,适合关注真实性能的研究者
我们系统研究了使用1,367小时广播语音及标准德语字幕对OpenAI Whisper large-v3进行微调以实现瑞士德语语音识别。在配备128 GB统一内存的NVIDIA DGX Spark(最高1 PFLOP FP4)上进行了16次迭代训练,对比了LoRA与全参数微调(1.55B参数),分析了幻觉成因,并量化了数据质量、字幕对齐度与训练策略的影响。最优模型在严格独立的瑞士德语方言测试集(ASGDTS)上测得25.6% WER;经风格差异分离的合理误差分析得内容词错误率(cWER)为13.8%,偏差校正后降至8.5%,表明真实错误率约为测量值的1/3。我们发现,现有最佳结果(17.1–17.5% WER)被基准污染夸大:未接触瑞士德语数据的Whisper自训练模型在测试集上已达13.88% WER,超越所有已发表系统。Phi-4-multimodal实验更显示3.9% WER,表明基准主要衡量惯例匹配而非方言理解。我们发布两个模型:LoRA适配器(25.32% WER,13.9% cWER)和全微调模型(25.60% WER,13.8% cWER),是少数公开可用、诚实评估的瑞士德语Whisper模型,遵循Apache 2.0许可,支持完全复现,无需机构数据协议。
原文摘要 · Abstract (English)
We present a systematic study of fine-tuning OpenAI's Whisper large-v3 for Swiss German ASR, using 1,367 hours of broadcast speech paired with Standard German subtitles as weak supervision. Through 16 iterative training runs on an NVIDIA DGX Spark (Grace Blackwell, 128 GB unified memory, up to 1 PFLOP FP4), we compare LoRA and full fine-tuning of the 1.55B-parameter model, investigate hallucination root causes, and quantify the effect of data quality, subtitle alignment, and training strategy. Our best model achieves 25.6% measured WER on the All Swiss German Dialects Test Set (ASGDTS) in an honest evaluation on strictly disjoint data. A harmonized error analysis separating genuine errors from valid stylistic variation (tense, word order, Swiss orthography) yields a content WER (cWER) of 13.8%, counting only actual recognition failures. Bias-corrected estimation reduces this to 8.5%, suggesting the true error rate is roughly one third of measured WER. We demonstrate that published state-of-the-art Swiss German ASR results (17.1-17.5% WER) are inflated by benchmark contamination: a vanilla Whisper model self-trained on the ASGDTS test set with zero Swiss German data achieves 13.88% WER, surpassing all published systems. Experiments with Phi-4-multimodal show an even stronger memorization effect (3.9% WER), revealing that the benchmark primarily measures convention matching rather than dialectal comprehension. We release two models, a LoRA adapter (25.32% WER, 13.9% cWER) and a full fine-tuned model (25.60% WER, 13.8% cWER), among the few publicly available, honestly evaluated Whisper models for Swiss German, under Apache 2.0 with full reproducibility, requiring no institutional data agreements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。