arXiv:2502.18215cs.CL2025-02

构建低资源语音平行语料库,助力小语种语音技术发展

Connecting Voices: LoReSpeech as a Low-Resource Speech Parallel Corpus

  • 基于协作平台与MFA工具,分步对齐短音频与长文本语音
  • 产出包含多语言对齐的LoReSpeech语料库,支持语音识别与翻译
  • 适合小语种研究、语音技术普惠及语言保护项目使用

对齐语音语料是自动语音识别(ASR)和语音翻译等自然语言处理技术的基础,但对未充分代表的语言而言仍极为稀缺,制约了其技术应用。本文提出一种构建低资源语音到语音翻译语料库LoReSpeech的方法。首先通过协作平台创建包含短音频及其转录的子语料库LoReASR;在此基础上,利用MFA等工具对圣经等长篇语音内容进行对齐。LoReSpeech提供跨语言与同语言的双重对齐,可推动多语言ASR系统、端到端语音翻译模型的发展,并助力语言保护与数字包容性建设。本工作属于Tutlayt AI项目(https://tutlayt.fr)。

原文摘要 · Abstract (English)

Aligned audio corpora are fundamental to NLP technologies such as ASR and speech translation, yet they remain scarce for underrepresented languages, hindering their technological integration. This paper introduces a methodology for constructing LoReSpeech, a low-resource speech-to-speech translation corpus. Our approach begins with LoReASR, a sub-corpus of short audios aligned with their transcriptions, created through a collaborative platform. Building on LoReASR, long-form audio recordings, such as biblical texts, are aligned using tools like the MFA. LoReSpeech delivers both intra- and inter-language alignments, enabling advancements in multilingual ASR systems, direct speech-to-speech translation models, and linguistic preservation efforts, while fostering digital inclusivity. This work is conducted within Tutlayt AI project (https://tutlayt.fr).

语音翻译低资源语料库语言保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。