对比CNN与Transformer在多语言语音质量评估中的表现差异
Language Barriers: Evaluating Cross-Lingual Performance of CNN and Transformer Architectures for Speech Quality Estimation
- 用CNN和Transformer模型在英语数据上训练,测试跨语言表现
- 普通话预测准确,瑞典语和荷兰语难建模,断续性最难预测
- 建议构建更均衡的多语言数据集以提升模型泛化能力
客观语音质量模型旨在通过自动化方法预测人耳感知的语音质量。然而,跨语言泛化仍是重大挑战,因为平均意见分(MOS)受语言、感知及数据集特性影响而存在差异。仅在英语数据上训练的模型难以适应具有不同音系、声调和语调特征的语言,导致评估结果不一致。本研究考察了两种语音质量模型:基于CNN的NISQA与基于Transformer的音频频谱变换器(AST)。两者均在超过49,000个英语语音样本上训练,随后在德语、法语、普通话、瑞典语和荷兰语上进行评估。使用皮尔逊相关系数(PCC)和均方根误差(RMSE)分析五项质量维度:失真、断续、响度、噪声及整体MOS。结果表明,尽管AST在跨语言表现上更稳定,但两类模型均存在明显偏差。普通话的质量预测与人类评分高度相关,而瑞典语和荷兰语则更具挑战性;所有语言中,断续性均难以建模。研究强调需构建更均衡的多语言数据集,并针对架构特性进行适配,以提升跨语言泛化能力。
原文摘要 · Abstract (English)
Objective speech quality models aim to predict human-perceived speech quality using automated methods. However, cross-lingual generalization remains a major challenge, as Mean Opinion Scores (MOS) vary across languages due to linguistic, perceptual, and dataset-specific differences. A model trained primarily on English data may struggle to generalize to languages with different phonetic, tonal, and prosodic characteristics, leading to inconsistencies in objective assessments. This study investigates the cross-lingual performance of two speech quality models: NISQA, a CNN-based model, and a Transformer-based Audio Spectrogram Transformer (AST) model. Both models were trained exclusively on English datasets containing over 49,000 speech samples and subsequently evaluated on speech in German, French, Mandarin, Swedish, and Dutch. We analyze model performance using Pearson Correlation Coefficient (PCC) and Root Mean Square Error (RMSE) across five speech quality dimensions: coloration, discontinuity, loudness, noise, and MOS. Our findings show that while AST achieves a more stable cross-lingual performance, both models exhibit noticeable biases. Notably, Mandarin speech quality predictions correlate highly with human MOS scores, whereas Swedish and Dutch present greater prediction challenges. Discontinuities remain difficult to model across all languages. These results highlight the need for more balanced multilingual datasets and architecture-specific adaptations to improve cross-lingual generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。