评测大模型在魁北克法语语音识别中的表现,发现通用数据集结果不适用本地场景。
Benchmarking Large Pretrained Multilingual Models on Québec French Speech Recognition
- 基于魁北克公共听证会录音构建本地评估基准
- 模型在魁北克法语上词错误率高于通用数据集表现
- 适合关注区域语言语音应用的开发者参考
我们评估了大型预训练多语言语音识别模型在加拿大魁北克地区法语方言上的表现,重点关注速度、词错误率和语义准确性。为此,我们基于CommissionsQC数据集构建了基准测试与评估流程,该数据集包含近期在魁北克举行的公共听证会中录制的自然对话。结果显示,这些模型在FLEURS或CommonVoice等知名基准上的公开性能,并不能准确预测其在CommissionsQC上的实际表现。研究结果对希望在真实场景或特定区域语言变体上构建语音应用的从业者具有重要参考价值。
原文摘要 · Abstract (English)
We evaluate the performance of large pretrained multilingual speech recognition models on a regional variety of French spoken in Québec, Canada, in terms of speed, word error rate and semantic accuracy. To this end we build a benchmark and evaluation pipeline based on the CommissionsQc datasets, a corpus of spontaneous conversations recorded during public inquiries recently held in Québec. Published results for these models on well-known benchmarks such as FLEURS or CommonVoice are not good predictors of the performance we observe on CommissionsQC. Our results should be of interest for practitioners interested in building speech applications for realistic conditions or regional language varieties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。