通过自适应混合后训练,提升语音对话模型的智能与表达力
WavAlign: Enhancing Intelligence and Expressiveness in Spoken Dialogue Models via Adaptive Hybrid Post-Training

- 分离语义与声学更新,仅在语义通道优化偏好
- 动态调节混合比例,避免不可靠梯度干扰
- 在多个基准上稳定提升语义质量与语音表现
端到端语音对话模型因具备更高表达力和感知潜力而受到关注,但当前开源模型的智能与表达力仍不理想。尽管在线强化学习在其他领域成功,直接迁移至语音对话仍面临挑战。本文从奖励建模与回溯采样角度分析障碍,发现稀疏偏好监督与共享参数下的密集语音生成存在冲突。为此提出一种模态感知的自适应后训练方法:将偏好更新限制在语义通道,通过显式锚定改善声学表现,并基于回溯统计动态调节两者混合比例,避免不可靠梯度。在多个语音对话基准与典型架构上评估,均实现语义质量与语音表达力的持续提升。
原文摘要 · Abstract (English)
End-to-end spoken dialogue models have garnered significant attention because they offer a higher potential ceiling in expressiveness and perceptual ability than cascaded systems. However, the intelligence and expressiveness of current open-source spoken dialogue models often remain below expectations. Motivated by the success of online reinforcement learning(RL) in other domains, one might attempt to directly apply preference optimization to spoken dialogue models, yet this transfer is non-trivial. We analyze these obstacles from the perspectives of reward modeling and rollout sampling, focusing on how sparse preference supervision interacts with dense speech generation under shared-parameter updates. Based on the analysis, we propose a modality-aware adaptive post-training recipe that makes RL practical for spoken dialogue: it constrains preference updates to the semantic channel and improves acoustic behavior via explicit anchoring, while dynamically regulating their mixture from rollout statistics to avoid unreliable preference gradients. We evaluate the method across multiple spoken dialogue benchmarks and representative architectures, and observe consistent improvements in semantic quality and speech expressiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。