arXiv:2608.31035cs.CLcs.SD2026-08

研究语音生成中感知预测器何时能有效作为强化学习奖励。

When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models

论文配图:When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models
图 1 · 摘自论文原文
  • 用GRPO优化多个感知维度的奖励,避免文本漂移。
  • 各感知奖励仅提升对应指标,不可互换使用。
  • 建议按预测轴分析奖励,适配多目标语音后训练。

基于编码器的文语转换(TTS)模型使语言模型后训练可用于语音生成,但尚不清楚学习到的感知预测器在何种情况下可作为强化学习奖励而不偏离人类听觉感知。本文通过组相对策略优化(GRPO)使用学习到的奖励来优化动漫风格说话、自然度、吸引力和唤醒度。为防止感知奖励因文本漂移而失真,引入字符错误率(CER)区域约束,并在相同奖励门控下对比了策略优化与Best-of-$N$重排序。单奖励实验显示,每项奖励仅显著提升对应指标,表明主观预测器并非可互换的质量代理。多评分者A/B测试显示人类迁移效果不均;奖励差距分析分离出平均迁移与轴内校准:总体分析中,符号化奖励差距显著预测听者选择,而残差CER差距无此效应,且轴内校准仍存在异质性。Best-of-8表现接近人类水平,且未明显劣于GRPO,说明GRPO应被视为将奖励选择行为摊入策略,而非普遍优于重排序。结果支持将主观语音奖励视为预测轴基元对,并为多奖励语音后训练前的奖励选择提供实用诊断方法。

原文摘要 · Abstract (English)

Codec-based text-to-speech (TTS) models make language-model post-training applicable to speech generation, but it remains unclear when learned perceptual predictors can serve as reinforcement learning rewards without losing alignment with human listeners. We study this question with Group Relative Policy Optimization (GRPO) using learned rewards for anime-like speaking style, naturalness, likability, and arousal. To prevent perceptual rewards from being optimized through transcript drift, we introduce a character error rate (CER) zone constraint and compare policy optimization with Best-of-$N$ reranking under the same reward gate. Across single-reward runs, each reward primarily improves its own target metric, showing that subjective predictors are not interchangeable quality surrogates. Multi-rater A/B tests further show uneven human transfer, while a reward-gap analysis separates average transfer from within-axis calibration: signed reward gaps significantly predict listener choices in the pooled analysis, whereas residual CER gaps do not, but per-axis calibration remains heterogeneous. Best-of-8 is a strong human-level baseline and is not clearly worse than GRPO perceptually, suggesting that GRPO should be viewed as amortizing reward-selected behavior into the policy rather than uniformly outperforming reranking. These results support analyzing subjective speech rewards as predictor-axis-base tuples and provide practical diagnostics for selecting rewards before multi-reward speech post-training.

语音生成强化学习感知评估奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。