用人类反馈强化学习优化语音增强,提升主观听感质量。
Using RLHF to align speech enhancement approaches to mean-opinion quality scores
- 基于MOS评分构建奖励模型,用RLHF微调语音增强模型。
- 在Voicebank+DEMAND数据集上,主观与客观指标均优于基线。
- 兼顾策略梯度与均方误差损失,实现多指标平衡优化。
客观语音质量评估指标常被用于评价语音增强算法,但研究表明它们作为学习目标时表现不佳,因与人类主观评分存在偏差,常导致明显失真和伪影,使增强效果失效。为解决此问题,我们提出一种基于人类反馈的强化学习(RLHF)框架,通过基于平均意见分(MOS)的奖励模型微调现有语音增强方法。实验结果表明,经RLHF微调的模型在Voicebank+DEMAND数据集上,对多种客观与基于MOS的语音质量评估指标均表现最优。消融实验证明,策略梯度损失与监督均方误差(MSE)损失均对跨指标的平衡优化至关重要。
原文摘要 · Abstract (English)
Objective speech quality measures are typically used to assess speech enhancement algorithms, but it has been shown that they are sub-optimal as learning objectives because they do not always align well with human subjective ratings. This misalignment often results in noticeable distortions and artifacts that cause speech enhancement to be ineffective. To address these issues, we propose a reinforcement learning from human feedback (RLHF) framework to fine-tune an existing speech enhancement approach by optimizing performance using a mean-opinion score (MOS)-based reward model. Our results show that the RLHF-finetuned model has the best performance across different benchmarks for both objective and MOS-based speech quality assessment metrics on the Voicebank+DEMAND dataset. Through ablation studies, we show that both policy gradient loss and supervised MSE loss are important for balanced optimization across the different metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。