用大模型生成真实场景下的手部动作,突破现有数据与技术瓶颈。
CLUTCH: Contextualized Language model for Unlocking Text-Conditioned Hand motion modelling in the wild
- 基于视觉-语言模型与3D追踪构建32K条野外手部动作数据集
- 提出新型分部位量化自编码器,提升动作重建精度与泛化能力
- 引入几何精修阶段,实现文本到动作的高质量对齐
手部动作在日常生活中至关重要,但自然场景下的手部动作建模仍鲜有研究。现有方法依赖于工作室采集的有限动作数据集,难以扩展至真实场景。此外,当前模型在文本与动作对齐及动画保真度方面表现不足。为此,本文(1)提出首个32,000条3D手部动作序列与对应文本的野外数据集3D-HIW,通过结合视觉-语言模型与先进3D手部追踪器构建大规模自注释流程;(2)提出基于大模型的CLUTCH系统,包含两项创新:(a)SHIFT——一种分部位模态分解的VQ-VAE架构,用于手部动作离散化;(b)几何精修阶段,通过直接作用于解码后手部参数的重建损失进行联合监督。实验表明,CLUTCH在文本到动作与动作到文本任务上均达到领先性能,建立了首个可扩展的真实场景手部动作建模基准。代码、数据与模型将公开。
原文摘要 · Abstract (English)
Hands play a central role in daily life, yet modeling natural hand motions remains underexplored. Existing methods that tackle text-to-hand-motion generation or hand animation captioning rely on studio-captured datasets with limited actions and contexts, making them costly to scale to "in-the-wild" settings. Further, contemporary models and their training schemes struggle to capture animation fidelity with text-motion alignment. To address this, we (1) introduce '3D Hands in the Wild' (3D-HIW), a dataset of 32K 3D hand-motion sequences and aligned text, and (2) propose CLUTCH, an LLM-based hand animation system with two critical innovations: (a) SHIFT, a novel VQ-VAE architecture to tokenize hand motion, and (b) a geometric refinement stage to finetune the LLM. To build 3D-HIW, we propose a data annotation pipeline that combines vision-language models (VLMs) and state-of-the-art 3D hand trackers, and apply it to a large corpus of egocentric action videos covering a wide range of scenarios. To fully capture motion in-the-wild, CLUTCH employs SHIFT, a part-modality decomposed VQ-VAE, which improves generalization and reconstruction fidelity. Finally, to improve animation quality, we introduce a geometric refinement stage, where CLUTCH is co-supervised with a reconstruction loss applied directly to decoded hand motion parameters. Experiments demonstrate state-of-the-art performance on text-to-motion and motion-to-text tasks, establishing the first benchmark for scalable in-the-wild hand motion modelling. Code, data and models will be released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。