用语音相似性补救语音搜索中冷门片名识别差的问题
Phonetically-Augmented Discriminative Rescoring for Voice Search Error Correction
- 基于ASR输出生成语音相近的候选词,拓展识别范围
- 融合语音候选与原识别结果,使错误率相对降低4.4%~7.6%
- 适合冷门影视名识别场景,提升语音搜索可用性
端到端(E2E)自动语音识别(ASR)模型依赖成对的音视频-文本数据训练,高质量标注需人工成本高。语音搜索应用(如数字媒体播放器)常因新片或小众片名在训练数据中缺失,导致识别效果差。本文提出一种语音增强的判别性重评分系统:首先基于ASR输出进行语音搜索,生成原系统未考虑的语音相似候选词;再通过重评分组件融合原始识别结果与语音候选,选择最终输出。实验表明,该方法在多个热门电影标题基准测试中,相对基线实现4.4%至7.6%的词错误率下降。
原文摘要 · Abstract (English)
End-to-end (E2E) Automatic Speech Recognition (ASR) models are trained using paired audio-text samples that are expensive to obtain, since high-quality ground-truth data requires human annotators. Voice search applications, such as digital media players, leverage ASR to allow users to search by voice as opposed to an on-screen keyboard. However, recent or infrequent movie titles may not be sufficiently represented in the E2E ASR system's training data, and hence, may suffer poor recognition. In this paper, we propose a phonetic correction system that consists of (a) a phonetic search based on the ASR model's output that generates phonetic alternatives that may not be considered by the E2E system, and (b) a rescorer component that combines the ASR model recognition and the phonetic alternatives, and select a final system output. We find that our approach improves word error rate between 4.4 and 7.6% relative on benchmarks of popular movie titles over a series of competitive baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。