用可解释的发音动作控制生成语音,让机器说话更像人。
Teaching Machines to Speak Using Articulatory Control
- 通过强化学习直接控制舌头、嘴唇等发音器官运动
- 生成语音与目标音节相似度超0.85,可被准确转录
- 适合研究语音生成机制或可解释人工智能的人
当前语音生成系统多依赖大型Transformer模型,缺乏可解释性且脱离人类发音的物理机制。本文提出一种新框架:通过显式发音动作控制生成语音,将语音生成视为类似机器人操作的运动控制任务。该方法利用强化学习训练策略,直接控制声道发音器官(如舌、唇、下颌)的运动以生成音节级语音。具体采用近端策略优化算法,基于音频感知器Sylber提供的声学反馈学习最优发音动作轨迹。生成的动作轨迹通过预训练的发音-语音解码器SPARC转化为音频。在六个目标音节上训练后,生成语音与目标音节的相似度超过0.85,对'please'、'loot'、'cat'等音节的人类转录准确率高,证明其可懂性。
原文摘要 · Abstract (English)
Current speech production systems predominantly rely on large transformer models that operate as black boxes, providing little interpretability or grounding in the physical mechanisms of human speech. We address this limitation by proposing a new framework: speech generation through explicit articulatory control. This reframes speech as a motor control task similar to robotic manipulation. Our approach uses reinforcement learning to train a policy that directly controls the movements of vocal tract articulators, such as the tongue, lips, and jaw, to produce syllable-level speech. Specifically, we employ the Proximal Policy Optimization algorithm to learn optimal articulatory movements based on acoustic feedback provided by our audio perceiver, Sylber. The resulting articulatory trajectories are decoded into audio using SPARC, a pre-trained articulatory-to-speech decoder. We train this framework on six target syllables, and it demonstrates successful convergence, with similarity scores between the policy-generated audio and the target syllables exceeding 0.85. Accurate human transcription of the audio for syllables such as "please", "loot", and "cat" demonstrates the intelligibility of this framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。