构建首个尼日利亚多口音语音翻译数据集,推动低资源语言研究
NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages
- 构建涵盖4种尼日利亚语言的50小时多口音语音对齐数据集
- 音频大模型在少样本下表现优于传统端到端和级联方法
- 为非洲低资源语言语音翻译提供可复现的基准与数据支持
低资源语言的语音翻译仍受限于高质量、多样化的平行语音数据稀缺,尤其在非洲语言环境中更为突出。为此,我们推出了NaijaS2ST,一个涵盖伊博语、豪萨语、约鲁巴语和尼日利亚皮钦语与英语配对的平行语音翻译数据集,每种语言约50小时语音,涵盖广泛说话人和口音,反映真实的多语言、多口音场景。基于该数据集,我们对级联、端到端(E2E)及音频大模型(AudioLLM)方法在双向翻译设置下进行了系统性基准测试。结果表明,使用少量示例的音频大模型在语音到文本翻译中表现优于微调后的级联与端到端方法;但在语音到语音翻译中,级联与音频大模型性能相当,说明针对此任务的专用模型仍有较大提升空间。通过提供高质量数据集与系统性基准,我们希望NaijaS2ST能成为推动低资源多语言语音翻译研究的重要基础。
原文摘要 · Abstract (English)
Speech translation for low-resource languages remains fundamentally limited by the scarcity of high-quality, diverse parallel speech data, a challenge that is especially pronounced in African linguistic contexts. To address this, we introduce NaijaS2ST, a parallel speech translation dataset spanning Igbo, Hausa, Yorùbá, and Nigerian Pidgin paired with English. The dataset comprises approximately 50 hours of speech per language and captures substantial variation in speakers and accents, reflecting realistic multilingual and multi-accent conditions. With NaijaS2ST, we conduct a comprehensive benchmark of cascaded, end-to-end (E2E), and AudioLLM-based approaches across bidirectional translation settings. Our results show that audio LLMs with few-shot examples are more effective for speech-to-text translation than cascaded and end-to-end methods trained on fine-tuned data. However, for speech-to-speech translation, the cascaded and audio LLM paradigms yield comparable performance, indicating that there is still considerable room for improvement in developing targeted, task-specific models for this setting. By providing both a high-quality dataset and a systematic benchmark, we hope that NaijaS2ST will serve as a strong foundation for advancing research in low-resource, multilingual speech translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。