提升语音合成中数字读法准确率,让机器念数更自然。
A Context-Based Numerical Format Prediction for a Text-To-Speech System
- 基于关键词、标点和符号提取数字上下文特征
- 分类准确率比现有方法提升30%至37%
- 适合需要精准数字发音的语音系统开发者
现有文本转语音系统在处理多种数字格式时往往表现不佳,导致合成语音可懂度下降。本文提出一种基于上下文的数字格式分类方法,可识别六类数字场景。通过提取关键词、标点和符号作为特征,结合支持向量机、K近邻、线性判别分析和决策树进行分类,并采用10折交叉验证评估性能。实验表明,该方法相较现有特征提取技术,在分类准确率上提升30%至37%。引入数字格式分类可显著提高语音合成系统的可懂度。
原文摘要 · Abstract (English)
Many of the existing TTS systems cannot accurately synthesize text containing a variety of numerical formats, resulting in reduced intelligibility of the synthesized speech. This research aims to develop a numerical format classifier that can classify six types of numeric contexts. Experiments were carried out using the proposed context-based feature extraction technique, which is focused on extracting keywords, punctuation marks, and symbols as the features of the numbers. Support Vector Machine, K-Nearest Neighbors Linear Discriminant Analysis, and Decision Tree were used as classifiers. We have used the 10-fold cross-validation technique to determine the classification accuracy in terms of recall and precision. It can be found that the proposed solution is better than the existing feature extraction technique with improvement to the classification accuracy by 30% to 37%. The use of the number format classification can increase the intelligibility of the TTS systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。