用物理模型生成同步语音,让发音动作与声音精准匹配。
Synchronized Three-Dimensional Vocal-Tract Motion for Speech Synchronization via Joint-Embedding Predictive Architecture Alignment
- 通过联合嵌入架构和强化学习优化发音动作序列
- 24个音素对测试中词错误率仅8.33%,自然度评分3.174
- 适合语音合成、人机交互与语音生理研究者
现代神经语音系统能生成可懂波形,但隐藏了发声的物理状态。相反,生物力学声道模型揭示了发音结构、接触行为、气流路径与几何约束,但直接物理波形合成仍不如现代神经声码器稳定。本文采用保持时长的声学载体生成听觉波形,同时用修正的三维声道模型生成同步的下颌、嘴唇、舌头、软腭、喉部、口腔气流和鼻腔气流运动。基于联合嵌入预测架构(JEPA)的表征与强化学习/交叉熵方法(RL/CEM)的轨迹选择循环,将发音动作对齐于声学载体及物理合理性约束。评估包含12组3D录音,覆盖24个最小对立音素对。在24词集上,声学载体取得8.33%的词错误率(WER)、4.17%的字符错误率(CER),UTMOS评分为3.174,平均JEPA得分0.864,平均音色保护得分0.947。
原文摘要 · Abstract (English)
Modern neural speech systems can generate intelligible waveforms, but they usually hide the physical speech-production state that produced the sound. Conversely, biomechanical vocal-tract models expose articulatory structure, contact behavior, airflow routing, and geometric constraints, but direct physical waveform synthesis remains less robust than modern neural vocoders. A duration-preserving acoustic carrier supplies the listening waveform, while a corrected three-dimensional vocal-tract model supplies synchronized jaw, lip, tongue, velum, laryngeal, oral-airflow, and nasal-airflow motion. A joint-embedding predictive architecture (JEPA)-style representation and a reinforcement learning/cross-entropy method (RL/CEM) trajectory-selection loop align articulatory actions to the acoustic carrier and to physical-plausibility constraints. The evaluation contains 12 3D recordings covering 24 minimal-pair stimuli. On the 24-word set, the carrier obtains good automatic speech recognition (ASR) results (an 8.33\% WER, a 4.17\% CER), a UTMOS score of 3.174, a mean JEPA score of 0.864, and a mean timbre-guard score of 0.947.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。