用音位属性提升手语动作生成自然度,效果优于现有方法。
Toward Phonology-Guided Sign Language Motion Generation: A Diffusion Baseline and Conditioning Analysis
- 采用扩散模型+SMPL-X表示,以音位特征为条件生成手语动作。
- 使用CLIP+自然语言转换的音位属性,性能超越SignAvatar所有指标。
- 音位符号转自然语言对CLIP有效,但对T5影响小,适合不同编码器。
基于文本输入生成自然、准确且视觉流畅的3D虚拟手语动作仍具挑战。本文训练了一个3D身体动作生成模型,探索音位属性(如手形、手位、动作)在手语生成中的作用,利用ASL-LEX 2.0标注数据。首先建立一个基于MDM风格扩散模型与SMPL-X表示的强基线,其在词素可辨识性指标上优于当前最先进的CVAE方法SignAvatar。随后系统研究不同文本编码器(CLIP vs. T5)、条件模式(仅词素 vs. 词素+音位属性)及属性表示形式(符号式 vs. 自然语言)的影响。结果表明:将符号式音位标注转化为自然语言是实现有效CLIP条件化的必要条件,而T5对此转化不敏感。最佳模型(CLIP+映射属性)在所有指标上均超越SignAvatar。该研究强调输入表示对文本编码器条件化的重要性,支持将词素与音位属性通过独立路径编码的结构化条件设计。
原文摘要 · Abstract (English)
Generating natural, correct, and visually smooth 3D avatar sign language motion conditioned on the text inputs continues to be very challenging. In this work, we train a generative model of 3D body motion and explore the role of phonological attribute conditioning for sign language motion generation, using ASL-LEX 2.0 annotations such as hand shape, hand location and movement. We first establish a strong diffusion baseline using an Human Motion MDM-style diffusion model with SMPL-X representation, which outperforms SignAvatar, a state-of-the-art CVAE method, on gloss discriminability metrics. We then systematically study the role of text conditioning using different text encoders (CLIP vs. T5), conditioning modes (gloss-only vs. gloss+phonological attributes), and attribute notation format (symbolic vs. natural language). Our analysis reveals that translating symbolic ASL-LEX notations to natural language is a necessary condition for effective CLIP-based attribute conditioning, while T5 is largely unaffected by this translation. Furthermore, our best-performing variant (CLIP with mapped attributes) outperforms SignAvatar across all metrics. These findings highlight input representation as a critical factor for text-encoder-based attribute conditioning, and motivate structured conditioning approaches where gloss and phonological attributes are encoded through independent pathways.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。