构建非洲语音技术基准,推动多语言包容性发展
Voice of a Continent: Mapping Africa's Speech Technology Frontier
- 建立非洲语音数据集与技术的系统性地图
- 推出SimbaBench基准与Simba系列模型,性能领先
- 揭示资源分布与语言家族对模型表现的影响
非洲丰富的语言多样性在语音技术中仍严重缺位,阻碍数字普惠。为此,我们系统性地梳理了非洲语音领域的数据集与技术现状,构建了新的综合性基准SimbaBench,用于下游非洲语音任务。基于SimbaBench,我们提出Simba系列模型,在多种非洲语言和语音任务上均实现先进性能。基准分析揭示了资源可用性的关键模式,模型评估表明数据质量、领域多样性和语言家族关系显著影响跨语言表现。本工作强调需拓展更反映非洲语言多样性的语音技术资源,并为未来更具包容性的语音技术研发奠定基础。
原文摘要 · Abstract (English)
Africa's rich linguistic diversity remains significantly underrepresented in speech technologies, creating barriers to digital inclusion. To alleviate this challenge, we systematically map the continent's speech space of datasets and technologies, leading to a new comprehensive benchmark SimbaBench for downstream African speech tasks. Using SimbaBench, we introduce the Simba family of models, achieving state-of-the-art performance across multiple African languages and speech tasks. Our benchmark analysis reveals critical patterns in resource availability, while our model evaluation demonstrates how dataset quality, domain diversity, and language family relationships influence performance across languages. Our work highlights the need for expanded speech technology resources that better reflect Africa's linguistic diversity and provides a solid foundation for future research and development efforts toward more inclusive speech technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。