arXiv:2510.02617cs.CV2025-10SIGGRAPH被引 2

用人体姿态引导注意力,实现实时语音同步视频生成

Input-Aware Sparse Attention for Real-Time Co-Speech Video Generation

  • 基于输入姿态关键点引导稀疏注意力,聚焦面部、手部等关键区域
  • 仅需少数去噪步骤即可实现1080p视频实时生成,速度提升3倍以上
  • 适合虚拟主播、智能客服等需要高帧率实时视频的应用场景

扩散模型可从音频生成逼真的语音同步视频,广泛应用于视频创作和虚拟人。然而,现有方法因需大量去噪步骤和高成本注意力机制而运行缓慢,难以实现实时部署。本文将多步扩散视频模型蒸馏为少步学生模型,但直接使用现有蒸馏方法会降低视频质量且无法满足实时性。为此,我们提出一种新蒸馏方法,利用输入人体姿态条件同时指导注意力和损失函数。首先,通过输入姿态关键点的精确对应关系,引导注意力关注说话者面部、手部和上半身等关键区域,减少冗余计算,增强身体部位的时间一致性,提升推理效率与动作连贯性。为进一步提升视觉质量,引入输入感知蒸馏损失,显著改善口型同步与手部动作真实感。结合输入感知稀疏注意力与蒸馏损失,本方法在保持实时性能的同时,相比近期音控与输入驱动方法实现了更优的视觉质量。大量实验验证了算法设计的有效性。

原文摘要 · Abstract (English)

Diffusion models can synthesize realistic co-speech video from audio for various applications, such as video creation and virtual agents. However, existing diffusion-based methods are slow due to numerous denoising steps and costly attention mechanisms, preventing real-time deployment. In this work, we distill a many-step diffusion video model into a few-step student model. Unfortunately, directly applying recent diffusion distillation methods degrades video quality and falls short of real-time performance. To address these issues, our new video distillation method leverages input human pose conditioning for both attention and loss functions. We first propose using accurate correspondence between input human pose keypoints to guide attention to relevant regions, such as the speaker's face, hands, and upper body. This input-aware sparse attention reduces redundant computations and strengthens temporal correspondences of body parts, improving inference efficiency and motion coherence. To further enhance visual quality, we introduce an input-aware distillation loss that improves lip synchronization and hand motion realism. By integrating our input-aware sparse attention and distillation loss, our method achieves real-time performance with improved visual quality compared to recent audio-driven and input-driven methods. We also conduct extensive experiments showing the effectiveness of our algorithmic design choices.

视频生成扩散模型实时生成姿态引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。