arXiv:2603.08125cs.CL2026-03被引 1

构建41小时阿联酋阿拉伯语语音库,助力低资源语言技术研究

Ramsa: A Large Sociolinguistically Rich Emirati Arabic Speech Corpus for ASR and TTS

  • 收录157名母语者录音,覆盖城市、游牧等方言及多元话题
  • 零样本测试下Whisper模型ASR错误率仅0.268,MMS-TTS表现最优
  • 专为社会语言学与语音技术研究设计,适合阿拉伯语低资源项目

Ramsa是一个正在建设的41小时阿联酋阿拉伯语语音语料库,旨在支持社会语言学研究和低资源语言技术发展。语料包含来自母语者结构化访谈和国家电视节目的录音片段,涵盖157名说话人(59名女性,98名男性),覆盖城市、游牧及山区/希希方言,并涉及文化遗产、农业与可持续性、日常生活、职业轨迹和建筑等主题。共包含91段独白和79段对话录音,长度与录制条件各异。使用10%子集评估商业与开源的自动语音识别(ASR)和文本转语音(TTS)模型在零样本设置下的性能,建立初步基线。Whisper-large-v3-turbo在ASR上表现最佳,平均词错误率和字符错误率分别为0.268和0.144;MMS-TTS-Ara在TTS上表现最佳,平均词错误率和字符错误率分别为0.285和0.081。这些基线具有竞争力,但仍存在较大提升空间。论文分析了面临的挑战并指明未来方向。

原文摘要 · Abstract (English)

Ramsa is a developing 41-hour speech corpus of Emirati Arabic designed to support sociolinguistic research and low-resource language technologies. It contains recordings from structured interviews with native speakers and episodes from national television shows. The corpus features 157 speakers (59 female, 98 male), spans subdialects such as Urban, Bedouin, and Mountain/Shihhi, and covers topics such as cultural heritage, agriculture and sustainability, daily life, professional trajectories, and architecture. It consists of 91 monologic and 79 dialogic recordings, varying in length and recording conditions. A 10\% subset was used to evaluate commercial and open-source models for automatic speech recognition (ASR) and text-to-speech (TTS) in a zero-shot setting to establish initial baselines. Whisper-large-v3-turbo achieved the best ASR performance, with average word and character error rates of 0.268 and 0.144, respectively. MMS-TTS-Ara reported the best mean word and character rates of 0.285 and 0.081, respectively, for TTS. These baselines are competitive but leave substantial room for improvement. The paper highlights the challenges encountered and provides directions for future work.

语音合成阿拉伯语低资源语言语料库

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。