用强化学习让语音模型精准定位并分类音频瑕疵。
Calibration-Reasoning Framework for Descriptive Speech Quality Assessment
- 分两阶段训练:先校准感知维度,再用特定奖励优化推理。
- 在QualiSpeech上达0.71均值皮尔逊相关系数,MOS预测提升13%。
- 适合需要可解释语音质量评估的工业场景与研究者。
可解释的语音质量评估需超越平均意见分(MOS),分析底层感知维度。为此,我们提出一种新型后训练方法,使基础音频大语言模型具备多维推理、检测和分类音频失真的能力。首先,通过校准阶段使模型对预定义的感知维度进行预测;其次,利用组相对策略优化(GRPO)结合维度特异性奖励,显著提升描述准确性与时序定位精度。该方法在多维QualiSpeech基准上达到0.71的均值皮尔逊相关系数(PCC),MOS预测性能提升13%。此外,细粒度的GRPO奖励大幅增强了模型对音频失真在时间上的精确定位与分类能力。
原文摘要 · Abstract (English)
Explainable speech quality assessment requires moving beyond Mean Opinion Scores (MOS) to analyze underlying perceptual dimensions. To address this, we introduce a novel post-training method that tailors the foundational Audio Large Language Model for multidimensional reasoning, detection and classification of audio artifacts. First, a calibration stage aligns the model to predict predefined perceptual dimensions. Second, a reinforcement learning stage leverages Group Relative Policy Optimization (GRPO) with dimension-specific rewards to heavily enhance accuracy of descriptions and temporal localization of quality issues. With this approach we reach state-of-the-art results of 0.71 mean PCC score on the multidimensional QualiSpeech benchmark and 13% improvement in MOS prediction driven by RL-based reasoning. Furthermore, our fine-grained GRPO rewards substantially advance the model's ability to pinpoint and classify audio artifacts in time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。