arXiv:2609.02291cs.CVcs.AI2026-09

用流模型实现一秒解码,视频压缩效率提升近六成

VoRTeC: Taming Foundation Flow for One-step Real time Video Compression

论文配图:VoRTeC: Taming Foundation Flow for One-step Real time Video Compression
图 1 · 摘自论文原文
  • 基于基础流模型压缩视频隐变量,预测压缩位置并融合多尺度先验
  • 相比扩散模型方法节省58%码率,解码速度提升3至197倍
  • 适合对实时性与画质要求高的视频传输场景

超低码率视频压缩仍面临严峻挑战:传统神经视频压缩易引入模糊伪影,而基于扩散的生成式压缩则存在解码延迟高、时间一致性差的问题。为此,我们提出基于基础流模型(Wan2.1)的视频压缩框架 $ t{VoRTeC}$。通过紧凑编码视频隐表示,预测压缩表示在流轨迹上的位置,并融合多尺度先验,$ t{VoRTeC}$ 能有效利用生成式视频流先验。无需访问流匹配网络的参数或梯度,即可实现一步解码和高感知保真重建。同时,通过尾帧复用和先验缓存,保持帧组间一致性。大量实验表明,本方法相比先前扩散类方法减少58%码率消耗,解码速度提升3至197倍:在720p下达13 FPS,480p下达32 FPS。

原文摘要 · Abstract (English)

Ultra-low bitrate video compression still faces critical challenges: traditional neural video compression inevitably introduces blurring artifacts, while diffusion-based generative video compression suffers from excessive decoding latency and poor temporal consistency. To address these issues, we propose $\mathtt{VoRTeC}$, a Video Compression framework built upon a foundational flow model (Wan2.1). By compactly encoding latent video representations, predicting the positions of compressed representations along flow trajectories, and integrating multi-scale priors, $\mathtt{VoRTeC}$ enables the compressor to harness generative video flow priors effectively. Without accessing the parameters or gradients of flow matching networks, our framework achieves one-step decoding and reconstructions with high perceptual fidelity. Meanwhile, we maintain consistency across frame groups via tail-frame reuse and prior caching. Extensive experiments demonstrate that our method reduces bit consumption by 58\% compared to prior diffusion-based approaches, with decoding speed boosted by 3 to 197 times: $\mathtt{VoRTeC}$ achieves a decoding speed of 13 FPS at 720p and 32 FPS at 480p.

视频压缩流模型实时解码低码率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。