研究法语语音识别中分词与自监督学习对效果的影响
A Comprehensive Analysis of Tokenization and Self-Supervised Learning in End-to-End Automatic Speech Recognition applied on French Language
- 对比不同子词分词算法和自监督模型的性能
- 使用多维度评估指标分析语音识别效果
- 适合关注语音识别细节优化的研究者
端到端自动语音识别(ASR)系统性能不断提升,已广泛应用于各类场景。尽管此类系统有诸多优势,但超参数与模型选择对其表现影响重大。传统上仅依赖字符错误率(CER)或词错误率(WER)评估,但多项研究表明这些指标不完整,难以准确反映下游应用需求。本文针对法语开展定性研究,从语言学与声学角度,系统考察子词分词算法与自监督学习模型的影响,采用综合评估指标进行分析。
原文摘要 · Abstract (English)
The performance of end-to-end automatic speech recognition (ASR) systems enables their increasing integration into numerous applications. While there are various benefits to such speech-to-text systems, the choice of hyperparameters and models plays a crucial role in their performance. Typically, these choices are determined by considering only the character (CER) and/or word error rate (WER) metrics. However, it has been shown in several studies that these metrics are largely incomplete and fail to adequately describe the downstream application of automatic transcripts. In this paper, we conduct a qualitative study on the French language that investigates the impact of subword tokenization algorithms and self-supervised learning models from different linguistic and acoustic perspectives, using a comprehensive set of evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。