首个公开的突尼斯阿拉伯语-英语混合语语音翻译数据集
TEDxTN: A Three-way Speech Translation Corpus for Code-Switched Tunisian Arabic - English
- 构建108场演讲的多语种语音数据集,含25小时带混用语言的口语
- 涵盖11个地区不同口音,支持语音识别与翻译基准测试
- 适合研究中东方言、多语言语音处理及低资源语言模型
本文提出TEDxTN,首个公开可用的突尼斯阿拉伯语至英语语音翻译数据集。为缓解阿拉伯方言数据稀缺问题,我们依据自定标注规范,收集、分割、转录并翻译了108场TEDx演讲,总计25小时语音,覆盖来自突尼斯11个不同地区的多种口音,包含大量语码转换现象。我们公开提供标注指南与数据集,便于后续扩展。同时报告了基于多个预训练和微调端到端模型的语音识别与语音翻译强基线结果。该数据集是首个开源且公开的代码切换突尼斯方言语音翻译语料库,有望推动突尼斯方言自然语言处理研究。
原文摘要 · Abstract (English)
In this paper, we introduce TEDxTN, the first publicly available Tunisian Arabic to English speech translation dataset. This work is in line with the ongoing effort to mitigate the data scarcity obstacle for a number of Arabic dialects. We collected, segmented, transcribed and translated 108 TEDx talks following our internally developed annotations guidelines. The collected talks represent 25 hours of speech with code-switching that cover speakers with various accents from over 11 different regions of Tunisia. We make the annotation guidelines and corpus publicly available. This will enable the extension of TEDxTN to new talks as they become available. We also report results for strong baseline systems of Speech Recognition and Speech Translation using multiple pre-trained and fine-tuned end-to-end models. This corpus is the first open source and publicly available speech translation corpus of Code-Switching Tunisian dialect. We believe that this is a valuable resource that can motivate and facilitate further research on the natural language processing of Tunisian Dialect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。