让AI根据个人特征生成匹配文本的个性化动作
PersonaBooth: Personalized Text-to-Motion Generation
- 用角色标记和多模态微调,让模型理解个人特征
- 在新数据集上实现更一致的角色化动作生成
- 适合做角色驱动动画、虚拟人交互的研究者
本文提出运动个性化任务,即基于包含角色特征的若干基础动作,生成与文本描述匹配的个性化动作。为此,我们构建了大规模运动数据集PerMo(PersonaMotion),捕捉多位演员的独特角色特征。针对预训练模型与角色数据分布差异大、动作内容多样难以保持角色一致性的问题,提出PersonaBooth方法:引入角色标记并进行文本与视觉的多模态适配;采用对比学习增强同角色样本内聚性;设计上下文感知融合机制,最大化多动作输入中的角色线索。实验表明,该方法显著优于现有运动风格迁移方法,建立了运动个性化的新基准。
原文摘要 · Abstract (English)
This paper introduces Motion Personalization, a new task that generates personalized motions aligned with text descriptions using several basic motions containing Persona. To support this novel task, we introduce a new large-scale motion dataset called PerMo (PersonaMotion), which captures the unique personas of multiple actors. We also propose a multi-modal finetuning method of a pretrained motion diffusion model called PersonaBooth. PersonaBooth addresses two main challenges: i) A significant distribution gap between the persona-focused PerMo dataset and the pretraining datasets, which lack persona-specific data, and ii) the difficulty of capturing a consistent persona from the motions vary in content (action type). To tackle the dataset distribution gap, we introduce a persona token to accept new persona features and perform multi-modal adaptation for both text and visuals during finetuning. To capture a consistent persona, we incorporate a contrastive learning technique to enhance intra-cohesion among samples with the same persona. Furthermore, we introduce a context-aware fusion mechanism to maximize the integration of persona cues from multiple input motions. PersonaBooth outperforms state-of-the-art motion style transfer methods, establishing a new benchmark for motion personalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。