将语音与人脸生成融合,实现音容同步的联合建模
UniTAF: A Modular Framework for Joint Text-to-Speech and Audio-to-Face Modeling
- 构建模块化框架,复用语音合成中间特征实现音脸联动
- 验证跨模态特征复用可行性,提升语音与表情一致性
- 适合研究多模态生成与交互系统设计的开发者参考
本文探讨将独立的文本转语音(TTS)与音频转人脸(A2F)模型整合为统一框架,以实现内部特征共享,从而提升由文本生成的语音与面部表情之间的一致性。同时,将情绪控制机制从TTS扩展至联合模型。本工作不以生成质量为目标,而是从系统设计角度验证了复用TTS中间表示进行语音与面部表情联合建模的可行性,并为后续语音表情协同设计提供工程实践参考。项目代码已开源:https://github.com/GoldenFishes/UniTAF
原文摘要 · Abstract (English)
This work considers merging two independent models, TTS and A2F, into a unified model to enable internal feature transfer, thereby improving the consistency between audio and facial expressions generated from text. We also discuss the extension of the emotion control mechanism from TTS to the joint model. This work does not aim to showcase generation quality; instead, from a system design perspective, it validates the feasibility of reusing intermediate representations from TTS for joint modeling of speech and facial expressions, and provides engineering practice references for subsequent speech expression co-design. The project code has been open source at: https://github.com/GoldenFishes/UniTAF
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。