arXiv:2601.16483eess.AS2026-01中稿 · ICASSP 2026被引 8

用在线强化学习提升语音增强模型的听感与任务表现。

FlowSE-GRPO: Training Flow Matching Speech Enhancement via Online Reinforcement Learning

  • 将在线GRPO算法适配到流匹配语音增强框架中
  • 仅用少量更新步骤即显著提升感知与任务指标
  • 多目标奖励策略有效避免音质下降,适合生成音频优化

生成式语音增强通过建模干净语音在噪声输入条件下的分布,为传统判别方法提供了替代方案。尽管强化学习(RL)在自然语言处理等领域已成功用于后训练对齐人类偏好与下游指标,但在语音增强领域仍以离线方法为主,尤其是在线方法研究较少。本文首次将在线组相对策略优化(GRPO)成功集成至流匹配语音增强框架,实现高效后训练对齐,仅需少量更新步骤即可提升感知与任务导向指标。不同于大语言模型中的GRPO应用,我们针对语音的连续时序特性及流匹配生成模型动态进行了算法适配。实验表明,单一奖励优化虽能快速提升得分,但常导致奖励滥用,降低音频保真度;为此,我们提出多指标奖励优化策略,在平衡不同目标的同时显著减少过拟合,全面提升性能。结果验证了在线GRPO在语音增强中的有效性,并为生成音频模型的强化学习后训练提供实用指导。

原文摘要 · Abstract (English)

Generative speech enhancement offers a promising alternative to traditional discriminative methods by modeling the distribution of clean speech conditioned on noisy inputs. Post-training alignment via reinforcement learning (RL) effectively aligns generative models with human preferences and downstream metrics in domains such as natural language processing, but its use in speech enhancement remains limited, especially for online RL. Prior work explores offline methods like Direct Preference Optimization (DPO); online methods such as Group Relative Policy Optimization (GRPO) remain largely uninvestigated. In this paper, we present the first successful integration of online GRPO into a flow-matching speech enhancement framework, enabling efficient post-training alignment to perceptual and task-oriented metrics with few update steps. Unlike prior GRPO work on Large Language Models, we adapt the algorithm to the continuous, time-series nature of speech and to the dynamics of flow-matching generative models. We show that optimizing a single reward yields rapid metric gains but often induces reward hacking that degrades audio fidelity despite higher scores. To mitigate this, we propose a multi-metric reward optimization strategy that balances competing objectives, substantially reducing overfitting and improving overall performance. Our experiments validate online GRPO for speech enhancement and provide practical guidance for RL-based post-training of generative audio models.

语音增强强化学习流匹配后训练优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。