MambaRate能准确评估不同采样率下的语音质量,减少采样率带来的偏差。
MambaRate: Speech Quality Assessment Across Different Sampling Rates
- 采用自监督嵌入与选择性状态空间建模,提升跨采样率的评估能力。
- 在少样本设置下比基线模型高14%,挑战赛中排名第四,仅落后冠军6%。
- 适用于高采样率语音质量评估,适合音频质量评测与模型开发人员参考。
我们提出MambaRate,用于预测语音波形的平均意见分(MOS),且对波形采样率的变化具有较低的偏差。该模型专为AudioMOS Challenge 2025的第3赛道设计,聚焦高采样频率语音的MOS预测。模型利用自监督嵌入和选择性状态空间建模,并通过高斯径向基函数(RBF)将目标评分编码为连续表示。挑战赛结果基于系统级斯皮尔曼等级相关系数(SRCC)进行评估。初始版本T16在无预训练的少样本设置下,相比预训练基线B03提升了约14%;在挑战赛中排名第四,与冠军系统差距约6%。此外,我们在BVCC数据集上展示额外结果,并对比不同输入表示的消融实验,新版本性能优于初始T16。
原文摘要 · Abstract (English)
We propose MambaRate, which predicts Mean Opinion Scores (MOS) with limited bias regarding the sampling rate of the waveform under evaluation. It is designed for Track 3 of the AudioMOS Challenge 2025, which focuses on predicting MOS for speech in high sampling frequencies. Our model leverages self-supervised embeddings and selective state space modeling. The target ratings are encoded in a continuous representation via Gaussian radial basis functions (RBF). The results of the challenge were based on the system-level Spearman's Rank Correllation Coefficient (SRCC) metric. An initial MambaRate version (T16 system) outperformed the pre-trained baseline (B03) by ~14% in a few-shot setting without pre-training. T16 ranked fourth out of five in the challenge, differing by ~6% from the winning system. We present additional results on the BVCC dataset as well as ablations with different representations as input, which outperform the initial T16 version.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。