arXiv:2503.08714cs.CVcs.AI2025-03中稿 · ACM MM2025被引 6

用语音和文字控制生成自然生动的人像说话视频

Versatile Multimodal Controls for Expressive Talking Human Animation

论文配图:Versatile Multimodal Controls for Expressive Talking Human Animation
图 1 · 摘自论文原文
  • 音频生成基础动作,文本控制具体行为,实现多模态灵活驱动
  • 支持从头部到全身的多尺度人物动画,生成表情丰富且语义准确的动作
  • 创新性地将3D动作令牌转为2D姿态,减少僵硬感,提升动作细节

在影视制作中,导演通常先让演员根据剧本自由表演,再提供具体动作指导。AI生成内容也面临类似需求:用户不仅需要基于音频自动生成口型同步和基础动作,还希望生成语义准确且富有表现力的全身动作,并能通过文本描述直接引导。为此,我们提出VersaAnimator,一个从任意肖像图生成富有表现力的说话人视频的通用框架。设计了运动生成器,从音频输入生成基础韵律动作,并支持文本提示控制特定动作。生成的全身3D动作令牌可适配不同尺度的人像,实现头部说话、半身手势乃至全身动作。此外,引入多模态可控视频扩散模型,以语音信号控制口型、面部表情和头部动作,同时以2D姿态引导身体动作。进一步设计了token2pose翻译器,将3D动作令牌平滑映射为2D姿态序列,缓解直接3D到2D转换带来的僵硬问题,增强动作细节。大量实验表明,VersaAnimator在保持身份一致性的前提下,实现了精准口型同步与富有表现力的语义化全身动作生成。

原文摘要 · Abstract (English)

In filmmaking, directors typically allow actors to perform freely based on the script before providing specific guidance on how to present key actions. AI-generated content faces similar requirements, where users not only need automatic generation of lip synchronization and basic gestures from audio input but also desire semantically accurate and expressive body movement that can be ``directly guided'' through text descriptions. Therefore, we present VersaAnimator, a versatile framework that synthesizes expressive talking human videos from arbitrary portrait images. Specifically, we design a motion generator that produces basic rhythmic movements from audio input and supports text-prompt control for specific actions. The generated whole-body 3D motion tokens can animate portraits of various scales, producing talking heads, half-body gestures and even leg movements for whole-body images. Besides, we introduce a multi-modal controlled video diffusion that generates photorealistic videos, where speech signals govern lip synchronization, facial expressions, and head motions while body movements are guided by the 2D poses. Furthermore, we introduce a token2pose translator to smoothly map 3D motion tokens to 2D pose sequences. This design mitigates the stiffness resulting from direct 3D to 2D conversion and enhances the details of the generated body movements. Extensive experiments shows that VersaAnimator synthesizes lip-synced and identity-preserving videos while generating expressive and semantically meaningful whole-body motions.

人物动画多模态控制视频生成动作迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。