用模仿学习提升老人听懂语音的能力,更自然且省数据。
Imitation Learning for Elder-Facing Speech Synthesis

- 通过专家示范学习老人友好的语音合成策略。
- 新方法在客观与主观测试中均优于传统监督与强化学习。
- 适合语音助手、老年健康科技等无障碍应用。
近年来文本转语音(TTS)技术已实现高度自然且富有表现力的语音生成,但现有系统主要针对成年人设计,忽视了老年人因年龄相关的感官与认知衰退带来的理解困难。以往研究通过收集老年人偏好反馈来调整模型参数,但获取足够偏好数据成本高、难度大,因老年人易疲劳。本文提出一种新型模仿学习(IL)框架,从专家示范中学习语音合成模型,并结合两阶段在线策略奖励学习(OPRL)改进群组相对策略优化(GRPO),以缓解有限专家示范下的奖励滥用问题。实验结果表明,采用OPRL的GRPO在客观与主观评估指标上均优于原始GRPO及监督基线。音频样例可于 https://dongru1.github.io/demo/im-efss 获取。
原文摘要 · Abstract (English)
Recent advances in text-to-speech (TTS) synthesis have achieved highly natural and expressive speech generation. However, these systems are designed for general adults and overlook older adults' speech comprehension needs due to age-related sensory and cognitive decline. Prior work involves older adults by collecting preference feedback to tune model parameters. However, obtaining sufficient preference data is costly and difficult, as older adults quickly become fatigued during collection. In this paper, we propose a novel imitation learning (IL) framework to learn TTS models from expert demonstrations. We further improve Group Relative Policy Optimization (GRPO) with two-stage on-policy reward learning (OPRL) to mitigate reward hacking under limited supervision from expert demonstration. Experimental results show that GRPO w/ OPRL outperforms GRPO and supervised baselines in objective and subjective metrics. Audio samples are available at https://dongru1.github.io/demo/im-efss
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。