模拟人耳听觉机制,提升语音合成质量评估准确性
APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech
- 基于生物听觉机理构建感知模块,将语音转为类电化学信号
- 引入残差向量量化建模语义失真,捕捉高层语音损伤
- 多模态融合架构结合渐进学习,更贴近人类主观判断
自动语音质量评估旨在通过计算模型量化人类对语音的主观感知,以减少人工评价的耗时。尽管基于深度学习的模型在预测平均意见分(MOS)方面已取得进展,但忽视基础听觉感知机制导致与人类判断一致性不足。为此,我们提出一种听觉感知引导的MOS预测模型(APG-MOS),通过融合听觉建模与语义分析,增强与人类判断的一致性。具体而言,首先设计一个基于生物听觉机制的感知模块,模拟耳蜗功能,将声学信号编码为生物对齐的电化学表示;其次提出基于残差向量量化(RVQ)的语义失真建模方法,量化语音在语义层面的退化;最后设计残差交叉注意力架构,并结合渐进学习策略,实现编码后的电化学信号与语义表示的多模态融合。实验表明,APG-MOS在两个主要基准上均表现优异。代码与检查点将在发表后公开于公共仓库。
原文摘要 · Abstract (English)
Automatic speech quality assessment aims to quantify subjective human perception of speech through computational models to reduce the need for labor-consuming manual evaluations. While models based on deep learning have achieved progress in predicting mean opinion scores (MOS) to assess synthetic speech, the neglect of fundamental auditory perception mechanisms limits consistency with human judgments. To address this issue, we propose an auditory perception guided-MOS prediction model (APG-MOS) that synergistically integrates auditory modeling with semantic analysis to enhance consistency with human judgments. Specifically, we first design a perceptual module, grounded in biological auditory mechanisms, to simulate cochlear functions, which encodes acoustic signals into biologically aligned electrochemical representations. Secondly, we propose a residual vector quantization (RVQ)-based semantic distortion modeling method to quantify the degradation of speech quality at the semantic level. Finally, we design a residual cross-attention architecture, coupled with a progressive learning strategy, to enable multimodal fusion of encoded electrochemical signals and semantic representations. Experiments demonstrate that APG-MOS achieves superior performance on two primary benchmarks. Our code and checkpoint will be available on a public repository upon publication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。