用人脸动画数据生成伪发音数据,提升语音转发音的精度
ArtBoost: Synthetic Articulatory Data Augmentation for Acoustic-to-Articulatory Inversion

- 从语音-人脸网格数据中提取伪发音轨迹用于预训练
- 在有限真实发音数据下,相关指标显著提升
- 适配多种模型架构,适合资源受限的发音研究
当前语音转发音(AAI)模型依赖昂贵且数据量有限的电磁发音仪(EMA)数据。为解决此问题,我们提出新型数据增强方法 ArtBoost,利用原本用于语音驱动3D人脸动画的大规模语音-网格数据集,通过可见面部关键点提取伪发音轨迹,先进行预训练,再在真实EMA数据上微调。实验表明,该方法在皮尔逊相关系数(PCC)和均方根误差(RMSE)上均有稳定提升。轨迹分析显示,伪发音信号反映了物理上合理的可见发音动态。跨不同AAI架构的额外评估也验证了性能提升的稳定性,表明ArtBoost可无缝集成至多种AAI模型中。结果说明,语音-网格数据是语音转发音任务中有效且可扩展的发音监督来源。
原文摘要 · Abstract (English)
Recent acoustic-to-articulatory inversion (AAI) models rely on electromagnetic articulography (EMA) data, which are costly and limited in scale. To address this limitation, we propose \textit{ArtBoost}, a novel data augmentation strategy that leverages large-scale speech--mesh datasets originally developed for speech-driven 3D facial animation to improve AAI under limited EMA supervision. \textit{ArtBoost} extracts pseudo articulatory trajectories from visible facial anchors and uses them for pre-training before fine-tuning on real EMA data. Experiments show consistent improvements in PCC and RMSE. Trajectory analyses confirm that the pseudo articulatory signals reflect physically meaningful visible articulatory dynamics. Additional evaluations across different AAI architectures demonstrate stable performance gains, indicating that \textit{ArtBoost} can be integrated into diverse AAI models. These results suggest that speech--mesh data provide an effective and scalable source of articulatory supervision for AAI. Project page: https://cau-irislab.github.io/Interspeech26-ArtBoost/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。