实现头身同步的实时风格化视频聊天肖像生成
ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model
- 分层运动扩散模型融合音视频输入,生成风格化表情与动作
- 支持512×768分辨率下30fps实时生成,含精细手势控制
- 适合需要自然肢体语言的虚拟主播、远程会议等场景
实时互动视频聊天肖像正成为未来趋势,尤其得益于文本与语音聊天技术的显著进步。现有方法多聚焦于头部动作的实时生成,但在身体动作与头部动作的同步性方面存在不足,且难以精细控制说话风格与面部细微表情。为此,我们提出一种新型框架,实现从说话头部到上半身交互的风格化实时肖像视频生成。该方法分为两阶段:第一阶段采用高效的分层运动扩散模型,基于音频输入综合显式与隐式运动表征,生成具有风格控制能力的多样化面部表情,并确保头部与身体动作同步;第二阶段生成包含上半身动作(如手势)的肖像视频,通过注入显式手部控制信号生成更精细的手势,并进行面部优化以提升整体真实感与表现力。本方法可在4090显卡上以最大512×768分辨率、最高30帧率实现高效连续生成,支持实时交互视频聊天。实验表明,该方法能生成富有表现力且自然的上半身动作视频。
原文摘要 · Abstract (English)
Real-time interactive video-chat portraits have been increasingly recognized as the future trend, particularly due to the remarkable progress made in text and voice chat technologies. However, existing methods primarily focus on real-time generation of head movements, but struggle to produce synchronized body motions that match these head actions. Additionally, achieving fine-grained control over the speaking style and nuances of facial expressions remains a challenge. To address these limitations, we introduce a novel framework for stylized real-time portrait video generation, enabling expressive and flexible video chat that extends from talking head to upper-body interaction. Our approach consists of the following two stages. The first stage involves efficient hierarchical motion diffusion models, that take both explicit and implicit motion representations into account based on audio inputs, which can generate a diverse range of facial expressions with stylistic control and synchronization between head and body movements. The second stage aims to generate portrait video featuring upper-body movements, including hand gestures. We inject explicit hand control signals into the generator to produce more detailed hand movements, and further perform face refinement to enhance the overall realism and expressiveness of the portrait video. Additionally, our approach supports efficient and continuous generation of upper-body portrait video in maximum 512 * 768 resolution at up to 30fps on 4090 GPU, supporting interactive video-chat in real-time. Experimental results demonstrate the capability of our approach to produce portrait videos with rich expressiveness and natural upper-body movements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。