实时生成高保真说话头像,速度超96帧/秒且不丢真。
SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads
- 用流式时序预训练+音频上下文缓存,稳住语音特征。
- 引入真实动作先验指导生成,避免长期误差累积。
- 适合需要低延迟互动的虚拟主播、数字人应用。
在音频驱动肖像生成中,实现高保真视觉质量与低延迟流式传输的平衡仍具挑战。现有大规模模型计算开销大,轻量级方案则常牺牲整体面部表征与时间稳定性。本文提出SoulX-FlashHead,一个统一的1.3B参数框架,支持实时、无限长度、高保真视频流生成。为应对流式场景下音频特征的不稳定性,引入流式感知时空预训练及时间音频上下文缓存机制,确保对短音频片段的鲁棒特征提取。为缓解长序列自回归生成中的误差累积与身份漂移问题,提出基于真实动作先验的双向蒸馏方法,提供精确物理引导。同时构建VividHead数据集,包含782小时严格对齐的高质量视频。大量实验表明,SoulX-FlashHead在HDTF与VFHQ基准上达到顶尖性能。其轻量版在单张NVIDIA RTX 4090上实现96 FPS推理速度,实现超快交互而不损失视觉连贯性。
原文摘要 · Abstract (English)
Achieving a balance between high-fidelity visual quality and low-latency streaming remains a formidable challenge in audio-driven portrait generation. Existing large-scale models often suffer from prohibitive computational costs, while lightweight alternatives typically compromise on holistic facial representations and temporal stability. In this paper, we propose SoulX-FlashHead, a unified 1.3B-parameter framework designed for real-time, infinite-length, and high-fidelity streaming video generation. To address the instability of audio features in streaming scenarios, we introduce Streaming-Aware Spatiotemporal Pre-training equipped with a Temporal Audio Context Cache mechanism, which ensures robust feature extraction from short audio fragments. Furthermore, to mitigate the error accumulation and identity drift inherent in long-sequence autoregressive generation, we propose Oracle-Guided Bidirectional Distillation, leveraging ground-truth motion priors to provide precise physical guidance. We also present VividHead, a large-scale, high-quality dataset containing 782 hours of strictly aligned footage to support robust training. Extensive experiments demonstrate that SoulX-FlashHead achieves state-of-the-art performance on HDTF and VFHQ benchmarks. Notably, our Lite variant achieves an inference speed of 96 FPS on a single NVIDIA RTX 4090, facilitating ultra-fast interaction without sacrificing visual coherence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。