用扩散模型生成多样手语数字人,同时保留表情与手势的语义细节。
Diverse Signer Avatars with Manual and Non-Manual Feature Modelling for Sign Language Production
- 基于潜空间扩散模型,从参考图生成逼真手语数字人。
- 在YouTube-SL-25数据集上,视觉质量显著优于现有方法。
- 显式建模面部与手势特征,支持跨族裔多样性表达。
手语表现的多样性对手语生成(SLP)至关重要,能体现外貌、面部表情和手部动作的差异。然而,现有模型难以在保持视觉质量的同时建模非手动特征(如情绪)。为此,我们提出一种新方法:利用潜空间扩散模型(LDM)从生成的参考图像合成逼真数字人,并设计了一种新的手语特征聚合模块,显式建模非手动特征(如面部)与手动特征(如手部)。实验表明,该模块在保留语言内容的前提下,可无缝使用不同族裔背景的参考图像实现多样性表达。在YouTube-SL-25数据集上的测试显示,本方法在感知评价指标上显著优于现有最先进方法,视觉质量更优。
原文摘要 · Abstract (English)
The diversity of sign representation is essential for Sign Language Production (SLP) as it captures variations in appearance, facial expressions, and hand movements. However, existing SLP models are often unable to capture diversity while preserving visual quality and modelling non-manual attributes such as emotions. To address this problem, we propose a novel approach that leverages Latent Diffusion Model (LDM) to synthesise photorealistic digital avatars from a generated reference image. We propose a novel sign feature aggregation module that explicitly models the non-manual features (\textit{e.g.}, the face) and the manual features (\textit{e.g.}, the hands). We show that our proposed module ensures the preservation of linguistic content while seamlessly using reference images with different ethnic backgrounds to ensure diversity. Experiments on the YouTube-SL-25 sign language dataset show that our pipeline achieves superior visual quality compared to state-of-the-art methods, with significant improvements on perceptual metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。