用奖励学习控制语音风格,零样本下自由调节语速与音高。
GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech

- 基于后生成奖励训练轻量级LoRA,解耦语音风格与说话人身份。
- 控制方向由语音长度和均值基频决定,错误率保障可懂性。
- 支持风格插值与多轴组合,无需重训练主模型。
我们提出GLASS,一种在零样本自回归语音合成中实现可组合声学风格控制的框架,其通过后生成奖励而非风格标签来学习控制。在零样本语音合成中,说话人提示常将说话人身份与语调特征(如语速、音高)纠缠,难以独立调整风格。GLASS将每个声学属性视为由奖励定义的控制方向:冻结语音合成主干,使用组相对策略优化(GRPO)训练一个轻量级LoRA适配器,以语音标记长度和均值基频作为风格奖励,以词错误率(WER)作为可懂性锚点。由于每个控制以LoRA权重更新形式表示,独立训练的适配器可通过线性LoRA运算交换、插值与组合,无需重训练主干。在语速与音高控制实验中,该方法实现了目标风格迁移,同时保持自然度、说话人相似性和可懂性,并展示了跨独立训练适配器的平滑插值与多轴组合能力。
原文摘要 · Abstract (English)
We propose GLASS, a framework for composable acoustic style control in zero-shot autoregressive text-to-speech (TTS) that learns controls from post-generation rewards rather than style labels. In zero-shot TTS, a speaker prompt often entangles speaker identity with prosodic attributes such as speaking rate and pitch, making it difficult to change style without changing the prompt itself. GLASS instead treats each acoustic attribute as a reward-defined control direction. For each control axis, GLASS freezes the TTS backbone and trains one lightweight LoRA adapter with Group Relative Policy Optimization (GRPO), using speech-token length and mean F0 as style rewards and WER as an intelligibility anchor. Because each control is represented as a LoRA weight update, independently trained adapters can be swapped, interpolated, and composed through linear LoRA arithmetic without retraining the backbone. Experiments on speaking rate and pitch control show targeted style shifts while preserving naturalness, speaker similarity, and intelligibility, and demonstrate smooth interpolation and multi-axis composition across independently trained adapters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。