arXiv:2409.15397eess.AScs.CL2024-09被引 9

用议会记录构建3种斯拉夫语的语音文本对齐数据集,超5000小时。

The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings

  • 基于议会文稿与录音,自动对齐长序列语音与文本
  • 完成克罗地亚、波兰、塞尔维亚语超5000小时高质量数据集
  • 方法可推广至更多低资源语言,助力语音技术发展

近年来语音与语言技术的进步既来自原始语言数据上的自监督方法,也得益于各类显式标注。为实现高质量语音处理,最有效的显式标注仍是语音信号与其对应文本转录的对齐,但这一数据在多数语言中仍缺失。本文提出一种基于议会会议文稿及录音,构建低资源语言语音-文本对齐数据集的方法。起点为26个欧洲国家议会的平行文稿语料库ParlaMint。在试点中,我们针对克罗地亚语、波兰语和塞尔维亚语,利用公开可得的录音扩展该语料库。主要挑战在于文稿与录音间缺乏全局对齐,且各模态数据顺序时常不一致,需在大搜索空间内对齐长序列语音与文本。最终成果为三个高质量数据集,覆盖超过5000小时语音与对应转录文本。尽管已显著提升三语言语音文本数据可用性,但该方法具有广泛推广潜力,适用于更多语言。

原文摘要 · Abstract (English)

Recent significant improvements in speech and language technologies come both from self-supervised approaches over raw language data as well as various types of explicit supervision. To ensure high-quality processing of spoken data, the most useful type of explicit supervision is still the alignment between the speech signal and its corresponding text transcript, which is a data type that is not available for many languages. In this paper, we present our approach to building large and open speech-and-text-aligned datasets of less-resourced languages based on transcripts of parliamentary proceedings and their recordings. Our starting point are the ParlaMint comparable corpora of transcripts of parliamentary proceedings of 26 national European parliaments. In the pilot run on expanding the ParlaMint corpora with aligned publicly available recordings, we focus on three Slavic languages, namely Croatian, Polish, and Serbian. The main challenge of our approach is the lack of any global alignment between the ParlaMint texts and the available recordings, as well as the sometimes varying data order in each of the modalities, which requires a novel approach in aligning long sequences of text and audio in a large search space. The results of this pilot run are three high-quality datasets that span more than 5,000 hours of speech and accompanying text transcripts. Although these datasets already make a huge difference in the availability of spoken and textual data for the three languages, we want to emphasize the potential of the presented approach in building similar datasets for many more languages.

语音对齐多语言数据集构建议会数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。