双分辨率注意力池化提升语音质量评分预测准确率
DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction
- 融合全局统计与局部显著段落的双重分析机制
- 在多个数据集上比平均池化提升10.39%相关性
- 适合语音质量评估、音频生成系统的性能验证
池化机制对均值意见分数(MOS)预测至关重要,能将可变长度音频特征转换为紧凑的固定尺寸表示,有效编码语音质量。现有方法通常仅采用单一粒度,聚焦全局或帧级分析,可能遗漏互补的感知信息。为此,我们提出双分辨率注意力统计池化(DRASP)框架,同时整合粗粒度全局统计摘要与细粒度感知显著段落的注意力分析。该双视角架构使模型能够更全面、稳健地表征语音质量,同时捕捉整体结构上下文与关键局部细节。大量实验验证了该框架的有效性与强泛化能力。其在多个数据集(MusicEval 和 AES-Natural)、多种MOS预测骨干网络(包括基于CLAP的模型和AudioBox-Aesthetics)及不同音频生成系统中,均持续优于各类基线方法,系统级斯皮尔曼等级相关系数(SRCC)相对平均池化提升10.39%。
原文摘要 · Abstract (English)
A pooling mechanism is essential for mean opinion score (MOS) prediction, facilitating the transformation of variable-length audio features into a concise fixed-size representation that effectively encodes speech quality. Existing pooling methods typically operate at a singular granularity, concentrating either on a comprehensive global perspective or a detailed frame-level analysis, which may overlook complementary perceptual insights. To address this limitation, we introduce the Dual-Resolution Attentive Statistics Pooling (DRASP) framework. DRASP integrates both coarse-grained, global statistical summaries and fine-grained, attentive analyses of perceptually significant segments. This dual-view architecture empowers our model to formulate a more thorough and robust representation, capturing both the overarching structural context and salient local details concurrently. Extensive experiments validate the effectiveness and strong generalization ability of the proposed framework. It consistently outperforms various baseline methods across diverse datasets (MusicEval and AES-Natural), MOS prediction backbones (including a CLAP-based model and AudioBox-Aesthetics), and different audio generation systems, achieving a relative improvement of 10.39% in system-level Spearman's rank correlation coefficient (SRCC) over the widely-used average pooling approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。