用多模态离散动作符生成手语,提升非手动特征表现。
M3T: Discrete Multi-Modal Motion Tokens for Sign Language Production
- 将人脸表情与身体模型融合,分模态量化生成动作符。
- 在三组基准上达最优,非手动特征区分任务准确率达58.3%。
- 适合手语生成、具身智能与无障碍交互研究者。
手语生成不仅需要手势动作,还需口型、眉毛抬升、视线方向和头部运动等非手动特征,这些在语法上不可或缺,无法仅从手部动作恢复。现有3D生成系统面临两大障碍:标准人体模型的面部空间维度过低,难以编码此类表达;而采用更丰富表示时,传统离散编码易出现码本坍塌,导致大部分表达空间不可达。本文提出SMPL-FX,将FLAME高维表情空间与SMPL-X身体模型结合,并采用针对身体、手部、面部的模态专用有限标量量化变分自编码器进行表征编码。M3T是基于此多模态动作符词汇训练的自回归变压器,辅以翻译目标促使嵌入具有语义基础。在三个标准基准(How2Sign、CSL-Daily、Phoenix14T)上,M3T均达到当前最优手语生成质量;在仅靠非手动特征区分的NMFs-CSL数据集上,准确率达到58.3%,优于最强对比基线的49.0%。
原文摘要 · Abstract (English)
Sign language production requires more than hand motion generation. Non-manual features, including mouthings, eyebrow raises, gaze, and head movements, are grammatically obligatory and cannot be recovered from manual articulators alone. Existing 3D production systems face two barriers to integrating them: the standard body model provides a facial space too low-dimensional to encode these articulations, and when richer representations are adopted, standard discrete tokenization suffers from codebook collapse, leaving most of the expression space unreachable. We propose SMPL-FX, which couples FLAME's rich expression space with the SMPL-X body, and tokenize the resulting representation with modality-specific Finite Scalar Quantization VAEs for body, hands, and face. M3T is an autoregressive transformer trained on this multi-modal motion vocabulary, with an auxiliary translation objective that encourages semantically grounded embeddings. Across three standard benchmarks (How2Sign, CSL-Daily, Phoenix14T) M3T achieves state-of-the-art sign language production quality, and on NMFs-CSL, where signs are distinguishable only by non-manual features, reaches 58.3% accuracy against 49.0% for the strongest comparable pose baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。