用专家混合模型提升语音质量评估,揭示层级差异根源
Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis
- 采用自监督模型+专家混合架构,适配多类语音评估任务
- 在话语级评估中表现仍受限,暴露现有方法深层瓶颈
- 提供跨层级评估的分析框架,适合语音系统开发者参考
自动语音质量评估在语音合成系统发展中至关重要,但现有模型在不同粒度预测任务中表现差异显著。本文提出基于自监督语音模型的增强型MOS预测系统,引入专家混合(MoE)分类头,并利用多个商用生成模型的合成数据进行数据增强。方法基于wav2vec2等自监督模型,设计专用MoE结构以应对多样化的语音质量评估任务。同时,构建了涵盖最新文本转语音、语音转换和语音增强系统的大型合成语音数据集。尽管采用MoE架构与扩展数据集,模型在句子级预测任务中的性能提升依然有限。研究揭示了当前方法在处理句子级质量评估时的局限性,为自动语音质量评估领域提供了新的技术路径,并深入探讨了不同评估粒度间性能差异的根本原因。
原文摘要 · Abstract (English)
Automatic speech quality assessment plays a crucial role in the development of speech synthesis systems, but existing models exhibit significant performance variations across different granularity levels of prediction tasks. This paper proposes an enhanced MOS prediction system based on self-supervised learning speech models, incorporating a Mixture of Experts (MoE) classification head and utilizing synthetic data from multiple commercial generation models for data augmentation. Our method builds upon existing self-supervised models such as wav2vec2, designing a specialized MoE architecture to address different types of speech quality assessment tasks. We also collected a large-scale synthetic speech dataset encompassing the latest text-to-speech, speech conversion, and speech enhancement systems. However, despite the adoption of the MoE architecture and expanded dataset, the model's performance improvements in sentence-level prediction tasks remain limited. Our work reveals the limitations of current methods in handling sentence-level quality assessment, provides new technical pathways for the field of automatic speech quality assessment, and also delves into the fundamental causes of performance differences across different assessment granularities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。