用大模型提升德语议会录音转写质量,构建801小时高质量语音文本数据集
Swiss Parliaments Corpus Re-Imagined (SPC_R): Enhanced Transcription with RAG-based Correction and Predicted BLEU
- 先用Whisper大模型转录,再用GPT-4o两次修正,重点纠错命名实体和语义不全段落
- 最终保留555小时高质量数据,相比原版提升6点BLEU,验证多阶段校正有效性
- 适合低资源领域语音研究者,尤其关注议会、法律等专业场景的语音处理应用
本文发布新版瑞士议会语料库(SPC_R),将完整的多小时苏黎世德语辩论会话(与官方会议记录对齐)转化为高质量语音-文本配对数据。管道流程首先在高算力条件下使用Whisper Large-v3将所有会话音频转为标准德语。随后通过两步GPT-4o修正:第一步结合原始转录与官方记录,修正主要为命名实体的误识别;第二步由独立GPT-4o评估每个段落的语义完整性。我们过滤掉预测BLEU分数(基于Whisper平均词元概率)和GPT-4o评分均低于阈值的段落。最终语料库包含801小时音频,其中555小时通过质量控制。相较于原句级SPC发布版本,本长格式数据集实现6点BLEU提升,证明结合强音转文字模型、大模型修正与数据驱动筛选,对低资源、特定领域语音语料具有显著增益。
原文摘要 · Abstract (English)
This paper presents a new long-form release of the Swiss Parliaments Corpus, converting entire multi-hour Swiss German debate sessions (each aligned with the official session protocols) into high-quality speech-text pairs. Our pipeline starts by transcribing all session audio into Standard German using Whisper Large-v3 under high-compute settings. We then apply a two-step GPT-4o correction process: first, GPT-4o ingests the raw Whisper output alongside the official protocols to refine misrecognitions, mainly named entities. Second, a separate GPT-4o pass evaluates each refined segment for semantic completeness. We filter out any segments whose Predicted BLEU score (derived from Whisper's average token log-probability) and GPT-4o evaluation score fall below a certain threshold. The final corpus contains 801 hours of audio, of which 555 hours pass our quality control. Compared to the original sentence-level SPC release, our long-form dataset achieves a 6-point BLEU improvement, demonstrating the power of combining robust ASR, LLM-based correction, and data-driven filtering for low-resource, domain-specific speech corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。