用统一框架实现人脸动画与视角控制,生成更自然的说话肖像。
STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits
- 通过软身份约束和视频多视角特性,实现2D域内的3D感知。
- 在多个基准上超越现有方法,生成动作更丰富、视角更连贯的视频。
- 适合需要高质量说话人图像生成的研究者与开发者。
本文提出STARCaster,一种身份感知的时空视频扩散模型,可在统一框架内实现语音驱动的人脸动画与动态视角控制,仅需身份嵌入或参考图像。现有2D语音-视频扩散模型严重依赖参考图,导致动作多样性不足;而3D感知方法通常依赖预训练三平面生成器进行反演,常引发重建不完美和身份漂移。本文从两个方面重构参考与几何建模范式:首先,在预训练阶段采用更宽松的身份约束;其次,通过视频数据固有的多视角特性,隐式实现3D感知。STARCaster采用分步设计:先构建身份感知的动作建模,再通过基于唇读的监督实现音画同步,最后借助时序到空间的转换完成新视角动画。为应对4D音视频数据稀缺问题,提出解耦学习策略,独立训练视角一致性与时间连贯性。大量实验表明,STARCaster在跨任务、跨身份场景下具有强泛化能力,在多个基准上持续优于先前方法。
原文摘要 · Abstract (English)
This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and dynamic viewpoint control, given an identity embedding or reference image, within a unified framework. Existing 2D speech-to-video diffusion models depend heavily on reference guidance, leading to limited motion diversity. At the same time, 3D-aware animation typically relies on inversion through pretrained tri-plane generators, which often leads to imperfect reconstructions and identity drift. We rethink reference- and geometry-based paradigms in two ways. First, we deviate from strict reference conditioning at pretraining by introducing softer identity constraints. Second, we address 3D awareness implicitly within the 2D video domain by leveraging the inherent multi-view nature of video data. STARCaster adopts a compositional approach progressing from ID-aware motion modeling, to audio-visual synchronization via lip reading-based supervision, and finally to novel view animation through temporal-to-spatial adaptation. To overcome the scarcity of 4D audio-visual data, we propose a decoupled learning approach in which view consistency and temporal coherence are trained independently. Comprehensive evaluations demonstrate that STARCaster generalizes effectively across tasks and identities, consistently surpassing prior approaches in different benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。