用教师学生框架从文本生成逼真手部动作,无需3D物体模型。
Teacher-Student Diffusion Model for Text-Driven 3D Hand Motion Generation
- 教师利用姿态参数提供结构化指导,学生仅凭文本学习动作生成。
- 在GRAB和H2O数据集上提升动作质量与多样性,效果稳定。
- 可适配多种辅助信号,推理时无需3D物体,通用性强。
从自然语言生成逼真的3D手部动作对虚拟现实、机器人和人机交互至关重要。现有方法或侧重全身动作而忽略精细手势,或需显式3D物体网格,限制了通用性。本文提出TSHaMo,一种模型无关的教师-学生扩散框架,用于文本驱动的手部动作生成。学生模型仅通过文本学习生成动作,教师则利用辅助信号(如MANO参数)在训练中提供结构化指导。协同训练策略使学生能借助教师的中间预测,推理时仍保持纯文本输入。在GRAB和H2O两个数据集上,采用两种扩散骨干网络进行评估,TSHaMo持续提升动作质量和多样性。消融实验验证其在使用多样化辅助输入时的鲁棒性与灵活性,且测试时无需3D物体。
原文摘要 · Abstract (English)
Generating realistic 3D hand motion from natural language is vital for VR, robotics, and human-computer interaction. Existing methods either focus on full-body motion, overlooking detailed hand gestures, or require explicit 3D object meshes, limiting generality. We propose TSHaMo, a model-agnostic teacher-student diffusion framework for text-driven hand motion generation. The student model learns to synthesize motions from text alone, while the teacher leverages auxiliary signals (e.g., MANO parameters) to provide structured guidance during training. A co-training strategy enables the student to benefit from the teacher's intermediate predictions while remaining text-only at inference. Evaluated using two diffusion backbones on GRAB and H2O, TSHaMo consistently improves motion quality and diversity. Ablations confirm its robustness and flexibility in using diverse auxiliary inputs without requiring 3D objects at test time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。