用自然语言生成手势视频,解决数据稀缺难题。
Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture Generation

- 基于提示的图像到视频模型生成真实感手势
- 合成手势与真人动作高度相似且更具多样性
- 适合动作识别、人机交互等领域的研究者
手势识别研究相比自然语言处理仍面临严重数据匮乏问题,进展受限于高昂的人类录制成本或无法生成真实多样性的图像处理方法。近期图像到视频基础模型的发展使得可由自然语言引导生成高保真、语义丰富的视频成为可能。本文提出并分析了基于提示的视频生成方法,构建了一个真实的指示性手势数据集,并对其下游任务效果进行严格评估。我们设计了一种数据生成流程,仅需少量人类参与者提供的参考样本即可生成指示性手势,为机器学习内外的研究者提供便捷路径。结果表明,合成手势在视觉保真度上与真实手势高度一致,同时引入有意义的变异性和新颖性,丰富了原始数据;使用混合数据集的多种深度模型性能更优。这些发现表明,即使处于早期阶段,图像到视频技术也能为手势合成提供强大的零样本方案,并显著提升下游任务表现。
原文摘要 · Abstract (English)
Gesture recognition research, unlike NLP, continues to face acute data scarcity, with progress constrained by the need for costly human recordings or image processing approaches that cannot generate authentic variability in the gestures themselves. Recent advancements in image-to-video foundation models have enabled the generation of photorealistic, semantically rich videos guided by natural language. These capabilities open up new possibilities for creating effort-free synthetic data, raising the critical question of whether video Generative AI models can augment and complement traditional human-generated gesture data. In this paper, we introduce and analyze prompt-based video generation to construct a realistic deictic gestures dataset and rigorously evaluate its effectiveness for downstream tasks. We propose a data generation pipeline that produces deictic gestures from a small number of reference samples collected from human participants, providing an accessible approach that can be leveraged both within and beyond the machine learning community. Our results demonstrate that the synthetic gestures not only align closely with real ones in terms of visual fidelity but also introduce meaningful variability and novelty that enrich the original data, further supported by superior performance of various deep models using a mixed dataset. These findings highlight that image-to-video techniques, even in their early stages, offer a powerful zero-shot approach to gesture synthesis with clear benefits for downstream tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。