构建全球方言与地区语言语音基准,评估模型表现并支持语音识别与生成优化。
Voxlect: A Speech Foundation Model Benchmark for Modeling Dialects and Regional Languages Around the Globe

- 基于30个公开语料库,覆盖20余种语言方言,构建多语言语音基准。
- 在超过200万语音样本上测试模型,发现方言分类结果符合地理连续性规律。
- 可应用于语音识别数据增强与语音生成系统评估,适合多语言研究者使用。
我们提出Voxlect,一个用于建模全球方言和地区语言的语音基础模型基准。涵盖英语、阿拉伯语、普通话、粤语、藏语、印地语系语言、泰语、西班牙语、法语、德语、巴西葡萄牙语和意大利语的方言与地区变体。研究使用来自30个公开语音语料库的超过200万条训练语音,均标注方言信息。评估了多个主流语音基础模型在方言分类上的性能,分析其在噪声条件下的鲁棒性,并通过错误分析揭示模型结果与地理分布的一致性。除基准评测外,还展示了Voxlect在下游应用中的潜力:可用于为现有语音识别数据集添加方言标签,实现对语音识别系统在方言差异下的性能细粒度分析;也可作为工具评估语音生成系统的表现。Voxlect已开源,许可遵循RAIL家族协议,地址:https://github.com/tiantiaf0627/voxlect。
原文摘要 · Abstract (English)
We present Voxlect, a novel benchmark for modeling dialects and regional languages worldwide using speech foundation models. Specifically, we report comprehensive benchmark evaluations on dialects and regional language varieties in English, Arabic, Mandarin and Cantonese, Tibetan, Indic languages, Thai, Spanish, French, German, Brazilian Portuguese, and Italian. Our study used over 2 million training utterances from 30 publicly available speech corpora that are provided with dialectal information. We evaluate the performance of several widely used speech foundation models in classifying speech dialects. We assess the robustness of the dialectal models under noisy conditions and present an error analysis that highlights modeling results aligned with geographic continuity. In addition to benchmarking dialect classification, we demonstrate several downstream applications enabled by Voxlect. Specifically, we show that Voxlect can be applied to augment existing speech recognition datasets with dialect information, enabling a more detailed analysis of ASR performance across dialectal variations. Voxlect is also used as a tool to evaluate the performance of speech generation systems. Voxlect is publicly available with the license of the RAIL family at: https://github.com/tiantiaf0627/voxlect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。