arXiv:2412.13609cs.CVcs.MM2024-12中稿 · AAAI被引 27

通过解耦肢体结构提升手语生成自然度

Sign-IDD: Iconicity Disentangled Diffusion for Sign Language Production

  • 将关节坐标拆分为骨骼方向与长度,更精准建模肢体关系
  • 在PHOENIX14T和USTC-CSL上生成效果优于现有方法
  • 适合关注手语生成真实感的研究者与开发者

手语生成(SLP)旨在从文本描述生成语义一致的手语视频,其中将文本词素转换为手语姿态(G2P)是关键步骤。现有方法通常将手语姿态视为离散的三维坐标并直接拟合,忽略了关节间的相对位置关系。为此,本文提出一种开创性的意象解耦扩散框架Sign-IDD,通过建模肢体骨骼来约束关节关联与手势细节,提升生成姿态的准确性和自然性。Sign-IDD引入新颖的意象解耦(ID)模块,将传统3D关节表示分解为4D骨骼表示,包含相邻关节间的3D空间方向向量与1D空间距离向量。同时,设计属性可控扩散(ACD)模块,通过属性分离层解耦骨骼方向与长度,并利用属性控制层引导姿态生成。ACD模块以词素嵌入作为语义条件,从噪声嵌入中生成手语姿态。在PHOENIX14T与USTC-CSL数据集上的大量实验验证了该方法的有效性。代码已公开于:https://github.com/NaVi-start/Sign-IDD。

原文摘要 · Abstract (English)

Sign Language Production (SLP) aims to generate semantically consistent sign videos from textual statements, where the conversion from textual glosses to sign poses (G2P) is a crucial step. Existing G2P methods typically treat sign poses as discrete three-dimensional coordinates and directly fit them, which overlooks the relative positional relationships among joints. To this end, we provide a new perspective, constraining joint associations and gesture details by modeling the limb bones to improve the accuracy and naturalness of the generated poses. In this work, we propose a pioneering iconicity disentangled diffusion framework, termed Sign-IDD, specifically designed for SLP. Sign-IDD incorporates a novel Iconicity Disentanglement (ID) module to bridge the gap between relative positions among joints. The ID module disentangles the conventional 3D joint representation into a 4D bone representation, comprising the 3D spatial direction vector and 1D spatial distance vector between adjacent joints. Additionally, an Attribute Controllable Diffusion (ACD) module is introduced to further constrain joint associations, in which the attribute separation layer aims to separate the bone direction and length attributes, and the attribute control layer is designed to guide the pose generation by leveraging the above attributes. The ACD module utilizes the gloss embeddings as semantic conditions and finally generates sign poses from noise embeddings. Extensive experiments on PHOENIX14T and USTC-CSL datasets validate the effectiveness of our method. The code is available at: https://github.com/NaVi-start/Sign-IDD.

手语生成扩散模型姿态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。