扩充匈牙利语对话语音识别数据集至200小时,提升模型训练效果。
Scaling Conversational Hungarian ASR: The BEA-Dialogue+ Corpus
- 放宽数据划分限制,保留主说话人分离,扩展至200小时对话数据
- 新数据集使无微调模型性能下降,但经SOT微调后词错率显著降低
- 适合研究匈牙利语对话语音识别的算法改进与评估
匈牙利语对话式语音识别受限于公开对话数据不足。BEA-Dialogue 数据集虽缓解此问题,但其严格区分说话人的训练/验证/测试集仅提供85小时可用数据。本文提出 BEA-Dialogue+,在保持主说话人完全分离的前提下,放宽实验者与对话伙伴的划分限制,最终获得200小时转录自然对话数据。该数据集支持对额外训练数据与说话人重叠之间权衡的受控研究。我们在两个数据集上评估了多个 Whisper 与 FastConformer 模型,包括基于 SOT(Serialized Output Training)的对话转录微调方法。结果表明,未经微调的模型在更大数据集上表现更差,而 SOT 微调能持续降低 WER、CER、cpWER 及 cpCER。整体而言,BEA-Dialogue+ 提供了一个更大且仍具挑战性的匈牙利语对话语音识别基准,是训练与评估对话转录系统的重要资源。
原文摘要 · Abstract (English)
Conversational automatic speech recognition in Hungarian is constrained by the limited amount of publicly available dialogue-style training data. The BEA-Dialogue corpus addresses this need, but its strictly speaker-disjoint train/dev/eval split reduces the usable material to only 85 hours. In this paper, we introduce BEA-Dialogue+, an expanded version of the corpus that relaxes the split criterion for experimenters and dialogue partners while preserving complete separation of the primary speakers. This results in 200 hours of transcribed natural conversations and enables a controlled study of the trade-off between additional training data and speaker overlap across the splits. We evaluate several Whisper- and FastConformer-based models on both corpus versions, including Serialized Output Training (SOT)-based fine-tuning for dialogue transcription. Our results show that the larger corpus is more challenging for models without fine-tuning, whereas SOT-based adaptation yields consistent improvements in WER, CER, cpWER, and cpCER. Overall, BEA-Dialogue+ provides a substantially larger yet still demanding benchmark for Hungarian dialogue ASR, and a practical resource for training and evaluating dialogue transcription systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。