arXiv:2606.21458cs.LG2026-06中稿 · Interspeech 2026

用感知奖励优化语音增强模型,提升听感质量。

Post-Training Speech Enhancement Language Models with Perceptual Rewards

论文配图:Post-Training Speech Enhancement Language Models with Perceptual Rewards
图 1 · 摘自论文原文
  • 后训练阶段用多指标感知奖励直接优化语音质量
  • 在DNS2020上达到当前最佳性能,显著提升听感
  • 多奖励融合避免单指标优化的奖励作弊问题

语音增强语言模型在离散音频标记上训练时表现优异,但其优化依赖于标记级交叉熵,而非评估中使用的感知指标。本文提出一种自回归语音增强语言模型的后训练方法,采用组序列策略优化(GSPO)并引入多指标感知奖励(DNSMOS、WER、UTMOS)。该方法直接以不可微的质量指标作为奖励信号,无需学习代理或离线偏好对。应用于两个基础模型UniSE和GenSE,本方法在DNS2020基准上取得最先进结果。人工评估消融实验表明,复合多指标奖励优于任一单一指标变体,证实多奖励优化可有效避免单指标训练中的奖励黑客问题。

原文摘要 · Abstract (English)

Speech enhancement language models achieve strong results when trained on discrete audio tokens, but their optimization relies on token-level cross-entropy rather than the perceptual metrics used for evaluation. We introduce a post-training stage for autoregressive speech enhancement language models using Group Sequence Policy Optimization (GSPO) with multi-metric perceptual rewards. Our method directly optimizes non-differentiable quality metrics (DNSMOS, WER, and UTMOS) as reward signals, without learned surrogates or offline preference pairs. Applied to two autoregressive base models, UniSE and GenSE, our approach achieves state-of-the-art results on the DNS2020 benchmark. A human evaluation ablation further shows that the composite multi-metric reward is preferred over any single-metric variant, confirming that multi-reward optimization avoids the reward hacking observed with single-metric training.

语音增强强化学习感知优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。