arXiv:2605.27062cs.CLcs.LG2026-05被引 1

构建了欧洲葡萄牙语议会发言的大规模带说话人标注语料库。

FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions

论文配图:FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions
图 1 · 摘自论文原文
  • 用先进语音识别模型对20年议会录音做转录对齐。
  • 含5800小时语音,4850小时有说话人标注,覆盖1180名议员。
  • 可显著提升欧洲葡萄牙语语音识别性能,相对错误率降14%。

当前最先进的自动语音识别(ASR)系统高度依赖大规模标注语料。由于欧洲葡萄牙语(EP)仅有约1100万使用者,远少于巴西葡萄牙语(约2亿),导致现有大规模语音数据资源中严重不足,影响了针对EP用户的语音系统性能。为弥补这一差距,我们构建了FalAR,一个大规模、带说话人标注的欧洲葡萄牙语议会会议语音语料库。该语料库覆盖约20年议会记录,总计5,800小时语音,其中4,850小时具备说话人身份标注,涵盖1,180位议员,附带年龄、性别、政治派别和议会职务等元数据。语料库通过先进的欧洲葡萄牙语CAMÕES ASR模型完成转录与对齐。本文详细描述数据采集流程及语料库主要特征,并评估数据量与对齐精度对ASR性能的影响。实验表明,将FalAR用于预训练可使基线模型的相对词错误率降低最高达14%。

原文摘要 · Abstract (English)

State-of-the-art performance for Automatic Speech Recognition (ASR) largely depends on the availability of large-scale labeled corpora. This creates a demand for increased data collection efforts, particularly for under-represented languages and dialectal varieties. Due to having considerably fewer speakers (around 11 million), European Portuguese (EP) is overshadowed by Brazilian Portuguese (BP) (around 200 million speakers) in currently available large-scale speech data resources, resulting in under-performing speech-based systems for EP users. To address this gap, and following similar data collection efforts for other languages, we present FalAR, a large-scale, speaker-annotated speech corpus of European Portuguese parliamentary sessions. Spanning approximately 20 years, FalAR comprises 5,800 hours of speech data. In addition, 4,850 hours have speaker identity annotations, for a total of 1,180 speakers with associated metadata including age, gender, political affiliation, and parliamentary role. The corpus was built using a state-of-the-art EP CAMÕES ASR model for transcription-reference alignment. In this paper, we describe the data collection process, together with the main characteristics of the FalAR corpus. Furthermore, we evaluate the trade-off between data quantity and alignment accuracy on ASR performance, with our experiments demonstrating that incorporating FalAR as pre-training data yields up to 14% relative WER improvement over baseline models.

语音识别语料库欧洲葡语说话人标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。