构建首个高质量突尼斯阿拉伯语语音与文本数据集,助力低资源方言语音识别
LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect
- 采集多种来源文本与真实场景语音,覆盖多说话人及阿英法混用场景
- 包含高精度语音转录,支持构建与评估突尼斯阿拉伯语语音识别系统
- 适合研究低资源语言语音识别、方言语音处理及跨语言混用建模者
由于突尼斯阿拉伯语的语系复杂性及标注语音数据集稀缺,开发该方言的自动语音识别(ASR)系统面临挑战。为此,我们提出 LinTO 音频与文本数据集——一套全面的资源,捕捉突尼斯阿拉伯语的音韵与词汇特征。数据集涵盖来自多个来源的多样化文本及真实世界音频样本,包含不同说话人以及突尼斯阿拉伯语与英语或法语之间的代码切换。通过提供高质量音频与精确转录,LinTO 数据集旨在为突尼斯阿拉伯语语音识别系统的构建与评测提供可靠材料。关键词:突尼斯阿拉伯语方言、语音转文本、低资源语言、音频数据增强
原文摘要 · Abstract (English)
Developing Automatic Speech Recognition (ASR) systems for Tunisian Arabic Dialect is challenging due to the dialect's linguistic complexity and the scarcity of annotated speech datasets. To address these challenges, we propose the LinTO audio and textual datasets -- comprehensive resources that capture phonological and lexical features of Tunisian Arabic Dialect. These datasets include a variety of texts from numerous sources and real-world audio samples featuring diverse speakers and code-switching between Tunisian Arabic Dialect and English or French. By providing high-quality audio paired with precise transcriptions, the LinTO audio and textual datasets aim to provide qualitative material to build and benchmark ASR systems for the Tunisian Arabic Dialect. Keywords -- Tunisian Arabic Dialect, Speech-to-Text, Low-Resource Languages, Audio Data Augmentation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。