arXiv:2607.10146eess.AS2026-07

对比三种模型在跨语料库语音质量评估中的表现,发现冻结的自监督特征更稳定可靠。

Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation

论文配图:Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation
图 1 · 摘自论文原文
  • 用冻结自监督特征+Transformer架构提升跨数据集泛化能力
  • 在英语纯净语料上,最佳模型达到MSE 0.36,接近顶尖水平
  • 结果表明该方案适合大规模通用语音质量评估场景

自动平均意见分(MOS)预测对大规模合成语音与音频增强系统评估至关重要,但模型常面临领域偏移问题。本研究全面评测三种架构:冻结自监督学习(SSL-FRZ)、微调自监督学习(SSL-FT)和视频视觉变压器(ViViT)。评估分两阶段进行:第一阶段使用包含19个多样化数据集的13万样本合并语料库;第二阶段聚焦17个纯英语数据集的净化语料库。采用系统性留一数据集外(LODO)协议,量化已见与未见分布间的泛化差距。最终,最优模型在ARECHO框架下与18种先进指标对比。结果表明,纯英语净化语料库在所有架构中均带来更高预测精度。虽然SSL-FT在已见数据上表现最优,但SSL-FRZ在未见分布中更具鲁棒性,在URGENT 2024基准上实现0.36的均方误差(MSE),接近领域优化的顶级模型(MSE 0.30)。尽管ViViT总容量低于基于SSL的模型,但在英语测试中表现稳定。LODO分析确认:模型在已见样本上表现显著更优,而结合冻结SSL嵌入与深度Transformer编码器的方案,是通用语音质量评估最稳定、可扩展的解决方案。为支持后续研究,最优的英语专用SSL-Transformer模型及权重已通过Hugging Face公开。

原文摘要 · Abstract (English)

Automatic Mean Opinion Score (MOS) prediction is essential for evaluating large-scale synthetic speech and audio enhancement systems, yet models frequently struggle with domain shift. This study presents a comprehensive benchmarking of three architectural frameworks: Frozen Self-Supervised Learning (SSL-FRZ), Fine-Tuned SSL (SSL-FT), and a Video Vision Transformer (ViViT). Evaluation is conducted in two phases: Part I utilizes a consolidated corpus of 130,000 samples across 19 diverse datasets, while Part II focuses on a purified 17-dataset English-only corpus. To assess robustness, a systematic Leave-One-Dataset-Out (LODO) protocol is employed to quantify the generalization gap between seen and unseen distributions. Finally, the top-performing model is benchmarked against 18 state-of-the-art (SOTA) metrics using the ARECHO framework. Results demonstrate that an English-only purified corpus consistently yields higher predictive precision across all architectures. While SSL-FT achieves the highest performance on seen validation data, the SSL-FRZ model provides superior robustness on unseen distributions, achieving a competitive Mean Squared Error (MSE) of 0.36 on the URGENT 2024 benchmark-closely matching domain-optimized SOTA metrics (MSE 0.30). Although the ViViT architecture remains below SSL-based models in total capacity, it delivers stable results in English-only trials. LODO results confirm that while models perform significantly better on seen samples, frozen SSL embeddings combined with deep Transformer encoders offer the most stable and scalable solution for universal speech quality assessment. To support further research, the top-performing English-only SSL-Transformer model and weights are made publicly available via Hugging Face.

语音评估自监督学习跨域泛化模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。