用听觉差异数据预训练语音质量评估模型,提升预测准确性。
JSQA: Speech Quality Assessment with Perceptually-Inspired Contrastive Pretraining Based on JND Audio Pairs
- 基于听觉可察觉差异对生成音频对,进行感知引导对比学习。
- 在NISQA数据集上,相比随机初始化模型,性能显著提升。
- 适合语音质量评估与感知建模研究者使用。
语音质量评估(SQA)旨在学习从高维输入到表示感知语音质量均值意见分数(MOS)的标量映射。由于感知差异和实验设计差异导致的MOS固有方差,此类映射学习极具挑战性。现有方法大多未将感知因素融入学习过程(仅依赖MOS标签),可能导致效果不佳。为此,我们提出JSQA,一种两阶段框架:首先利用听觉可察觉差异(JND)音频对进行感知引导的对比预训练,随后在NISQA数据集上微调以预测MOS。JND对由干净的LibriSpeech语句与来自CHiME-3的背景噪声在不同信噪比(SNRs)下混合生成。预训练后的编码器在相同网络结构下微调后,在多种评估指标上表现显著优于从零开始训练的模型。结果表明,将感知因素融入预训练过程能极大提升SQA模型性能。
原文摘要 · Abstract (English)
Speech quality assessment (SQA) is often used to learn a mapping from a high-dimensional input space to a scalar that represents the mean opinion score (MOS) of the perceptual speech quality. Learning such a mapping is challenging for many reasons, but largely because MOS exhibits high levels of inherent variance due to perceptual and experimental-design differences. Many solutions have been proposed, but many approaches do not properly incorporate perceptual factors into their learning algorithms (beyond the MOS label), which could lead to unsatisfactory results. To this end, we propose JSQA, a two-stage framework that pretrains an audio encoder using perceptually-guided contrastive learning on just noticeable difference (JND) pairs, followed by fine-tuning for MOS prediction. We first generate pairs of audio data within JND levels, which are then used to pretrain an encoder to leverage perceptual quality similarity information and map it into an embedding space. The JND pairs come from clean LibriSpeech utterances that are mixed with background noise from CHiME-3, at different signal-to-noise ratios (SNRs). The encoder is later fine-tuned with audio samples from the NISQA dataset for MOS prediction. Experimental results suggest that perceptually-inspired contrastive pretraining significantly improves the model performance evaluated by various metrics when compared against the same network trained from scratch without pretraining. These findings suggest that incorporating perceptual factors into pretraining greatly contributes to the improvement in performance for SQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。