构建6000万词斯洛伐克议会语料库,提升低资源语音识别性能
SloPal: A 60-Million-Word Slovak Parliamentary Corpus with Aligned Speech and Fine-Tuned ASR Models
- 用锚点对齐技术构建2806小时音视频对齐数据集
- 微调Whisper使错误率降低70%,小模型逼近大模型效果
- 开源完整语料与模型,助力斯洛伐克语语音研究
斯洛伐克语仍属低资源语音识别语言,公开训练数据不足100小时。本文提出SloPal,一个涵盖2001–2024年、包含33万段发言者分割文本(6600万词,2.2亿词元)的综合性斯洛伐克议会语料库,附带发言人姓名、职务和会议信息等丰富元数据。从中构建出SloPalSpeech,一个2806小时的音视频对齐数据集,每段最长30秒,采用无语言依赖锚点对齐流程生成,专为Whisper语音识别模型训练优化。在SloPalSpeech上微调Whisper可使词错误率(WER)降低高达70%;其小型模型(244M参数)性能接近大型基础模型(15亿参数),仅需六分之一参数量。我们已将SloPal文本语料、对齐音频及四个微调后的Whisper模型公开发布于https://huggingface.co/collections/NaiveNeuron/slopal,为当前最全面的开放斯洛伐克议会语言资源。
原文摘要 · Abstract (English)
Slovak remains a low-resource language for automatic speech recognition (ASR), with fewer than 100 hours of publicly available training data. We present SloPal, a comprehensive Slovak parliamentary corpus comprising 330,000 speaker-segmented transcripts (66 million words, 220 million tokens) spanning 2001--2024, with rich metadata including speaker names, roles, and session information. From this collection, we derive SloPalSpeech, a 2,806-hour aligned speech dataset with segments up to 30 seconds, constructed using a language-agnostic anchor-based alignment pipeline and optimized for Whisper-based ASR training. Fine-tuning Whisper on SloPalSpeech reduces Word Error Rate (WER) by up to 70\%, with the fine-tuned small model (244M parameters) approaching base large-v3 (1.5B parameters) performance at 6$\times$ fewer parameters. We publicly release the SloPal text corpus, SloPalSpeech aligned audio, and four fine-tuned Whisper models at https://huggingface.co/collections/NaiveNeuron/slopal, providing the most comprehensive open Slovak parliamentary language resource to date.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。