构建儿童语音纠错数据集,提升儿童语音识别准确率
CHSER: A Dataset and Case Study on Generative Speech Error Correction for Child ASR
- 构建包含20万对的儿童语音纠错数据集CHSER
- 在零样本设置下相对降低28.5%的错误率
- 适合关注儿童语音识别与纠错的研究者
自动语音识别(ASR)系统在处理儿童语音时表现不佳,主要因其声学和语言特征差异大,且儿童语音数据集稀缺,导致转录错误率高。尽管成人语音纠错方法已取得进展,但其在儿童语音上的效果仍不清楚。为此,我们提出CHSER,一个面向生成式语音纠错(GenSEC)的儿童语音纠错数据集,包含20万条跨年龄和语风的假设-转录对。实验表明,在零样本设置下微调可实现最高28.5%的相对单词错误率(WER)下降;在微调过的ASR系统上应用则可降低13.3%。此外,误差分析显示,生成式纠错虽能改善替换和删除错误,但在插入错误和儿童特有的不流畅表达上仍存在困难。这些发现凸显了生成式纠错在提升儿童语音识别方面的潜力。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) systems struggle with child speech due to its distinct acoustic and linguistic variability and limited availability of child speech datasets, leading to high transcription error rates. While ASR error correction (AEC) methods have improved adult speech transcription, their effectiveness on child speech remains largely unexplored. To address this, we introduce CHSER, a Generative Speech Error Correction (GenSEC) dataset for child speech, comprising 200K hypothesis-transcription pairs spanning diverse age groups and speaking styles. Results demonstrate that fine-tuning on the CHSER dataset achieves up to a 28.5% relative WER reduction in a zero-shot setting and a 13.3% reduction when applied to fine-tuned ASR systems. Additionally, our error analysis reveals that while GenSEC improves substitution and deletion errors, it struggles with insertions and child-specific disfluencies. These findings highlight the potential of GenSEC for improving child ASR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。