首个面向欧洲葡萄牙语的语音识别评测基准,填补资源空白。
CAMÕES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese
- 构建涵盖46小时多领域数据的评测基准与425小时训练集
- 微调模型与自训练模型表现相当,相对零样本模型提升超35%错误率
- 适合研究欧洲葡语语音识别或跨语言迁移的研究者
现有葡萄牙语自动语音识别资源主要聚焦于巴西葡语,导致欧洲葡萄牙语(EP)及其他变体研究不足。为弥补这一缺口,我们提出CAMÕES,首个针对欧洲葡语及其他葡语变体的开放框架。该框架包含:(1) 覆盖46小时欧洲葡语测试数据的综合性评测基准,涵盖多个领域;(2) 一系列先进模型。我们评估了多种基础模型的零样本与微调性能,以及从头训练的E-Branchformer模型。使用425小时精选的欧洲葡语数据进行微调与训练。结果表明,微调基础模型与E-Branchformer在欧洲葡语上表现相当。最佳模型相比最强零样本模型,相对词错误率(WER)降低超过35%,确立了欧洲葡语及其他变体的新基准。
原文摘要 · Abstract (English)
Existing resources for Automatic Speech Recognition in Portuguese are mostly focused on Brazilian Portuguese, leaving European Portuguese (EP) and other varieties under-explored. To bridge this gap, we introduce CAMÕES, the first open framework for EP and other Portuguese varieties. It consists of (1) a comprehensive evaluation benchmark, including 46h of EP test data spanning multiple domains; and (2) a collection of state-of-the-art models. For the latter, we consider multiple foundation models, evaluating their zero-shot and fine-tuned performances, as well as E-Branchformer models trained from scratch. A curated set of 425h of EP was used for both fine-tuning and training. Our results show comparable performance for EP between fine-tuned foundation models and the E-Branchformer. Furthermore, the best-performing models achieve relative improvements above 35% WER, compared to the strongest zero-shot foundation model, establishing a new state-of-the-art for EP and other varieties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。