arXiv:2604.23632cs.CVcs.MM2026-04被引 9

实时生成音视频虚拟形象,速度提升99倍且保持高质量。

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation

论文配图:Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
图 1 · 摘自论文原文
  • 异步双流扩散架构+未来语音线索预测,减少生成延迟。
  • 20.38帧/秒运行,延迟仅0.94秒,比基线快16倍以上。
  • 适合需要高保真音视频同步的实时交互场景。

实时文本驱动的音视频虚拟形象生成需同步合成高保真视频与语音,但现有音频-视觉扩散模型速度过慢,加速后质量显著下降。我们提出Hallo-Live,一种结合异步双流扩散与以人为本偏好引导蒸馏的流式框架。为降低因果生成中的口型延迟,引入未来扩展注意力(Future-Expanding Attention),使每帧视频可访问同步音频及短期未来语音线索。为缓解少步蒸馏带来的质量损失,提出人本偏好引导的动态模型蒸馏(HP-DMD),基于视觉保真度、语音自然度与音画同步性重加权训练样本。在两块NVIDIA H200 GPU上,系统实现20.38 FPS、0.94秒延迟,吞吐量提升16.0倍,延迟降低99.3倍,相较教师模型Ovi表现更优。尽管加速显著,仍保持相近的VideoAlign与Sync Confidence得分,在整体质量-效率权衡中优于其他加速基线。定性结果表明其在写实、多说话人及风格化场景下均有稳健泛化能力。据我们所知,Hallo-Live是首个结合流式双流扩散与偏好引导蒸馏的实时文本驱动音视频生成框架。

原文摘要 · Abstract (English)

Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio-visual diffusion models remain too slow for interactive use and often degrade noticeably after aggressive acceleration. We present Hallo-Live, a streaming framework for joint audio-visual avatar generation that combines asynchronous dual-stream diffusion with human-centric preference-guided distillation. To reduce articulation lag in causal generation, we introduce Future-Expanding Attention, which allows each video block to access synchronous audio together with a short horizon of future phonetic cues. To mitigate the quality loss of few-step distillation, we further propose Human-Centric Preference-Guided DMD (HP-DMD), which reweights training samples using rewards from visual fidelity, speech naturalness, and audio-visual synchronization. On two NVIDIA H200 GPUs, Hallo-Live runs at 20.38 FPS with 0.94 seconds latency, yielding 16.0x higher throughput and 99.3x lower latency than the teacher model Ovi. Despite this speedup, it retains strong generation quality, reaching comparable VideoAlign overall score and Sync Confidence score while outperforming other accelerated baselines in the overall quality-efficiency trade-off. Qualitative results further show robust generalization across photorealistic, multi-speaker, and stylized scenarios. To the best of our knowledge, Hallo-Live is the first framework to combine streaming dual-stream diffusion with preference-guided distillation for real-time, text-driven audio-visual generation.

音视频生成实时渲染扩散模型虚拟形象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。