14B模型实现实时无限音视频驱动虚拟人生成,延迟低于1秒。
SoulX-FlashTalk: Real-Time Infinite Streaming of Audio-Driven Avatars via Self-Correcting Bidirectional Distillation
- 用双向蒸馏保留时空关联,提升动作连贯性与细节。
- 自纠正机制避免长期生成误差累积,支持无限流式输出。
- 全栈优化实现32帧/秒实时运行,启动延迟仅0.87秒。
将大规模扩散模型用于实时、无限时长的音视频驱动虚拟人生成面临巨大工程挑战,主要源于计算负载与严格延迟限制之间的矛盾。现有方法常通过强制单向注意力或降低模型容量来妥协视觉质量。为此,我们提出SoulX-FlashTalk,一个专为高保真实时流式生成优化的140亿参数框架。不同于传统单向范式,采用自纠正双向蒸馏策略,在视频块内保留双向注意力,有效维持关键时空相关性,显著提升动作连贯性与视觉细节。为保障无限生成稳定性,引入多步回溯自纠正机制,使模型可自主恢复累积误差,防止崩溃。此外,构建了包含混合序列并行、并行VAE及内核级优化的全栈推理加速套件。大量评估表明,SoulX-FlashTalk是首个实现140亿规模、子秒级启动延迟(0.87秒)并达到32帧/秒实时吞吐的系统,为高保真交互式数字人合成树立新标准。
原文摘要 · Abstract (English)
Deploying massive diffusion models for real-time, infinite-duration, audio-driven avatar generation presents a significant engineering challenge, primarily due to the conflict between computational load and strict latency constraints. Existing approaches often compromise visual fidelity by enforcing strictly unidirectional attention mechanisms or reducing model capacity. To address this problem, we introduce \textbf{SoulX-FlashTalk}, a 14B-parameter framework optimized for high-fidelity real-time streaming. Diverging from conventional unidirectional paradigms, we use a \textbf{Self-correcting Bidirectional Distillation} strategy that retains bidirectional attention within video chunks. This design preserves critical spatiotemporal correlations, significantly enhancing motion coherence and visual detail. To ensure stability during infinite generation, we incorporate a \textbf{Multi-step Retrospective Self-Correction Mechanism}, enabling the model to autonomously recover from accumulated errors and preventing collapse. Furthermore, we engineered a full-stack inference acceleration suite incorporating hybrid sequence parallelism, Parallel VAE, and kernel-level optimizations. Extensive evaluations confirm that SoulX-FlashTalk is the first 14B-scale system to achieve a \textbf{sub-second start-up latency (0.87s)} while reaching a real-time throughput of \textbf{32 FPS}, setting a new standard for high-fidelity interactive digital human synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。