arXiv:2608.01978cs.CV2026-08

用情绪代理+低秩缓存,实现实时一人一拍情感可控人脸动画

Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation

论文配图:Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation
图 1 · 摘自论文原文
  • 构建情绪驱动的代理头像,一次训练可生成多表情动作视频
  • 引入低秩缓存机制,推理时减少90%以上显存占用
  • 适合做短视频、虚拟主播等需要快速生成情感化人脸内容的场景

基于扩散模型的音频驱动人脸动画发展迅速,但实现实时单张图像、情感可控的人脸动画仍具挑战。现有方法常因缺乏情感感知的动作先验,以及多步去噪过程中的高成本外观计算而受限。为此,我们提出代理头像结合低秩缓存的级联框架,实现高效实时的一次性情感可控人脸动画。该方法不直接从音频生成目标人脸,而是使用基于高斯分布的情绪代理头像作为可复用的动作生成器,仅需对单一身份训练一次,即可根据音频与情感标签生成富有表现力的驱动视频。由于代理头像仅提供运动信息而不包含目标外观或几何,因此通过大规模一次性重定向模型提取与身份无关的运动特征,并适配至任意目标人脸。为提升推理效率,引入零样本外观复用的低秩缓存机制,在初始去噪步骤缓存参考外观特征,并用轻量级低秩适配器建模后续特征变化。大量实验表明,本方法在情感表现力、身份保持性和推理开销方面均显著优于现有方法,实现了真正意义上的实时一次性人脸动画。

原文摘要 · Abstract (English)

Audio-driven portrait animation has advanced rapidly with diffusion-based generative models, yet real-time one-shot generation with expressive emotion control remains challenging. Existing methods often suffer from insufficient emotion-aware motion priors and expensive appearance computation during multi-step denoising. To address these issues, we propose Proxy Avatar Meets Low-Rank Caching, a cascaded framework for real-time one-shot emotion-controllable portrait animation. Instead of directly generating the target portrait from audio, our method uses a Gaussian-based emotion proxy avatar as a reusable motion generator, which is trained once on a single identity to produce expressive driving videos from audio and emotion labels. Since the proxy avatar only provides motion rather than target appearance or geometry, a large-scale one-shot retargeting model further extracts identity-independent motion from the proxy performance and adapts it to arbitrary target portraits. To improve inference efficiency, we introduce zero-shot appearance reuse with low-rank caching, which caches reference appearance features at the initial denoising step and models subsequent feature variations using lightweight low-rank adapters. Extensive experiments demonstrate that our method achieves stronger emotional expressiveness, better identity-preserving animation, and substantially reduced inference cost, enabling real-time one-shot portrait animation.

人脸动画情感控制实时生成低秩缓存

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。