用自然语言生成3D手部动作,支持手语和非手语场景。
Text-Driven 3D Hand Motion Generation from Sign Language Data
- 基于大尺度手语数据与LLM自动构建文本-动作配对数据
- 模型在跨手语、跨动作类型下保持稳定生成效果
- 适合手语生成、人机交互与虚拟角色动画研究者
本研究旨在训练一个以自然语言描述为条件的3D手部动作生成模型,描述内容包括手势形状、位置及手指/手/臂的运动特征。为此,我们利用大规模手语视频数据集,结合噪声伪标注的手语类别,通过大语言模型(LLM)结合手语属性词典与运动脚本提示,自动生成3D手部动作与对应文本描述的配对数据,规模前所未有。该数据支持训练出文本条件的手部动作扩散模型HandMDM,具备跨领域鲁棒性:不仅可生成同一手语中未见的手势类别,还可泛化至其他手语及非手语手部动作。我们进行了全面的实验验证,并将公开训练好的模型与数据集,以推动该新兴领域的后续研究。
原文摘要 · Abstract (English)
Our goal is to train a generative model of 3D hand motions, conditioned on natural language descriptions specifying motion characteristics such as handshapes, locations, finger/hand/arm movements. To this end, we automatically build pairs of 3D hand motions and their associated textual labels with unprecedented scale. Specifically, we leverage a large-scale sign language video dataset, along with noisy pseudo-annotated sign categories, which we translate into hand motion descriptions via an LLM that utilizes a dictionary of sign attributes, as well as our complementary motion-script cues. This data enables training a text-conditioned hand motion diffusion model HandMDM, that is robust across domains such as unseen sign categories from the same sign language, but also signs from another sign language and non-sign hand movements. We contribute extensive experimental investigation of these scenarios and will make our trained models and data publicly available to support future research in this relatively new field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。