用分块令牌压缩视频,超低码率下仍保高清
TVC: Tokenized Video Compression with Ultra-Low Bit Rate
- 双流架构:离散与连续令牌分路处理
- 离散流经掩码+无损压缩,连续流用上下文模型量化
- 适合超低码率视频传输,如卫星、物联网场景
令牌化视觉表示在图像压缩中表现优异,但其向视频领域的拓展受限于复杂的时序动态和严格的码率约束。本文提出令牌化视频压缩(TVC),一种基于令牌的双流框架,可在超低码率下有效运行。TVC采用Cosmos视频令牌化器提取离散与连续令牌流。离散令牌通过策略性掩码后,利用离散棋盘上下文模型无损压缩以降低传输开销;掩码后的令牌由仅解码器的Transformer通过时空令牌预测重建。同时,连续令牌经量化后,使用连续棋盘上下文模型压缩,提供互补的连续信息。解码端通过基于ControlNet的多尺度融合模块整合两路信号,确保高感知质量与稳定重建保真度。本工作验证了令牌化视频压缩的可行性,并指明语义感知、原生令牌化方法的新方向。
原文摘要 · Abstract (English)
Tokenized visual representations have shown promise in image compression, yet their extension to video remains underexplored due to the challenges posed by complex temporal dynamics and stringent bit rate constraints. In this paper, we present tokenized video compression (TVC), a token-based dual-stream framework designed to operate effectively at ultra-low bit rates. TVC leverages the Cosmos video tokenizer to extract both discrete and continuous token streams. The discrete tokens are partially masked using a strategic masking scheme and then compressed losslessly with a discrete checkerboard context model to reduce transmission overhead. The masked tokens are reconstructed by a decoder-only Transformer with spatiotemporal token prediction. In parallel, the continuous tokens are quantized and compressed using a continuous checkerboard context model, providing complementary continuous information at ultra-low bit rates. At the decoder side, the two streams are fused with a ControlNet-based multi-scale integration module, ensuring high perceptual quality alongside stable fidelity in reconstruction. Overall, this work illustrates the practicality of tokenized video compression and points to new directions for semantics-aware, token-native approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。