分析低资源语音识别中转录不一致的影响,发现非主要挑战
Investigating Transcription Normalization in the Faetar ASR Benchmark
- 用人工构建词典分析转录问题,定位核心难点
- 发现词典约束解码有帮助,但二元语言模型无益
- 适合研究低资源语音识别与数据质量的学者
我们研究了在费塔尔(Faetar)自动语音识别基准中,转录不一致性的作用。该基准属于极具挑战性的低资源语音识别任务。借助一个小型人工构建的词典,我们发现尽管转录中确实存在不一致性,但这并非任务的主要难点。同时,我们证明基于二元语法的词级语言模型并未带来额外收益,而将解码过程限制在有限词典内则有一定帮助。该任务仍然极为困难。
原文摘要 · Abstract (English)
We examine the role of transcription inconsistencies in the Faetar Automatic Speech Recognition benchmark, a challenging low-resource ASR benchmark. With the help of a small, hand-constructed lexicon, we conclude that find that, while inconsistencies do exist in the transcriptions, they are not the main challenge in the task. We also demonstrate that bigram word-based language modelling is of no added benefit, but that constraining decoding to a finite lexicon can be beneficial. The task remains extremely difficult.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。