融合语义与声学特征,提升语音自然度评分预测精度
SAMOS: A Neural MOS Prediction Model Leveraging Semantic Representations and Acoustic Features
- 用wav2vec2提取语义表征,用BiVocoder提取声学特征
- 在BVCC数据集上超越现有最佳模型,BC2019上表现相当
- 适合语音合成系统自动评估与质量优化研究者
使用均值意见分(MOS)预测模型评估语音自然度对语音合成系统的自动评估具有重要意义。早期的MOS预测模型仅以原始波形或幅度谱为输入,而较先进的方法则利用自监督学习(SSL)模型从语音中提取语义表征进行预测。但这些方法仅使用语音信息的有限方面,导致预测精度受限。为此,本文提出SAMOS,一种同时利用语音语义与声学信息进行MOS预测的模型。具体而言,SAMOS采用预训练的wav2vec2提取语义表征,并使用预训练BiVocoder的特征提取器获取声学特征。这两种特征随后输入包含多任务头和聚合层的预测网络,以获得最终的MOS分数。实验结果表明,所提SAMOS在BVCC数据集上的系统级评估指标优于当前最先进的MOS预测模型,在BC2019数据集上表现相当。
原文摘要 · Abstract (English)
Assessing the naturalness of speech using mean opinion score (MOS) prediction models has positive implications for the automatic evaluation of speech synthesis systems. Early MOS prediction models took the raw waveform or amplitude spectrum of speech as input, whereas more advanced methods employed self-supervised-learning (SSL) based models to extract semantic representations from speech for MOS prediction. These methods utilized limited aspects of speech information for MOS prediction, resulting in restricted prediction accuracy. Therefore, in this paper, we propose SAMOS, a MOS prediction model that leverages both Semantic and Acoustic information of speech to be assessed. Specifically, the proposed SAMOS leverages a pretrained wav2vec2 to extract semantic representations and uses the feature extractor of a pretrained BiVocoder to extract acoustic features. These two types of features are then fed into the prediction network, which includes multi-task heads and an aggregation layer, to obtain the final MOS score. Experimental results demonstrate that the proposed SAMOS outperforms current state-of-the-art MOS prediction models on the BVCC dataset and performs comparable performance on the BC2019 dataset, according to the results of system-level evaluation metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。