解决多模态视频生成延迟高、流式推理不稳定问题
EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation

- 分阶段知识迁移+自适应校准,实现单步推理
- 长时序训练中仅对尾帧强制对齐,提升稳定性
- 修复高频细节,适合实时视频生成场景
近期多模态视频生成模型虽具备高视觉质量,但其高昂延迟与有限时间稳定性阻碍了实时部署。流式推理进一步加剧问题,导致空间模糊、时间漂移和唇音不同步等多模态退化,形成难以调和的效率-性能权衡。为此,我们提出EchoTorrent,一种四重设计的新架构:(1) 多教师训练通过在不同偏好领域微调预训练模型,获得特定领域专家,并依次将领域知识传递给学生模型;(2) 自适应CFG校准(ACC-DMD)通过分阶段时空调度校正音频CFG增强误差,在去冗余计算的同时支持每步单次推理;(3) 混合长尾强制机制,在长时序自回放训练中仅对尾帧强制对齐,结合因果-双向混合结构,有效缓解流式模式下的时空退化并增强参考帧保真度;(4) VAE解码器精炼器通过像素域优化恢复高频细节,避免潜在空间歧义。大量实验与分析表明,EchoTorrent实现了少步自回归生成,显著延长时间一致性、身份保留与音唇同步能力。
原文摘要 · Abstract (English)
Recent multi-modal video generation models have achieved high visual quality, but their prohibitive latency and limited temporal stability hinder real-time deployment. Streaming inference exacerbates these issues, leading to pronounced multimodal degradation, such as spatial blurring, temporal drift, and lip desynchronization, which creates an unresolved efficiency-performance trade-off. To this end, we propose EchoTorrent, a novel schema with a fourfold design: (1) Multi-Teacher Training fine-tunes a pre-trained model on distinct preference domains to obtain specialized domain experts, which sequentially transfer domain-specific knowledge to a student model; (2) Adaptive CFG Calibration (ACC-DMD), which calibrates the audio CFG augmentation errors in DMD via a phased spatiotemporal schedule, eliminating redundant CFG computations and enabling single-pass inference per step; (3) Hybrid Long Tail Forcing, which enforces alignment exclusively on tail frames during long-horizon self-rollout training via a causal-bidirectional hybrid architecture, effectively mitigates spatiotemporal degradation in streaming mode while enhancing fidelity to reference frames; and (4) VAE Decoder Refiner through pixel-domain optimization of the VAE decoder to recover high-frequency details while circumventing latent-space ambiguities. Extensive experiments and analysis demonstrate that EchoTorrent achieves few-pass autoregressive generation with substantially extended temporal consistency, identity preservation, and audio-lip synchronization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。