一模型同时生成自然表情与手势,减少参数量且效果更优
Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters
- 用适配器共享跨模态特征,统一生成面部与手势动作
- 在两个任务上均达当前最佳性能,参数量显著降低
- 适合需要高效协同生成人脸与肢体动作的场景
近期在伴随说话手势和逼真说话人脸生成方面进展显著,但多数方法仅聚焦单一任务。尝试同时生成两者的模型常依赖独立模块,增加训练复杂度并忽视面部与身体动作的内在关联。为此,本文提出一种新模型架构,在单个网络中联合生成面部与身体运动。该方法通过适配器实现跨模态共享权重,将不同模态映射到共同潜在空间。实验表明,所提框架不仅保持了当前最优的说话手势与真实人脸生成性能,还显著减少了模型所需参数量。
原文摘要 · Abstract (English)
Recent advances in co-speech gesture and talking head generation have been impressive, yet most methods focus on only one of the two tasks. Those that attempt to generate both often rely on separate models or network modules, increasing training complexity and ignoring the inherent relationship between face and body movements. To address the challenges, in this paper, we propose a novel model architecture that jointly generates face and body motions within a single network. This approach leverages shared weights between modalities, facilitated by adapters that enable adaptation to a common latent space. Our experiments demonstrate that the proposed framework not only maintains state-of-the-art co-speech gesture and talking head generation performance but also significantly reduces the number of parameters required.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。