用生成模型实现0.02%超低视频压缩,接收端靠算力恢复画质
Generative Video Compression: Towards 0.01% Compression Rate for Video Transmission
- 将视频编码为极小表示,由接收端生成模型重建画面
- 实测压缩率达0.02%,在特定条件下突破0.01%阈值
- 适合带宽受限场景,可在消费级GPU快速推理
能否实现低至0.01%的视频压缩率?我们提出生成式视频压缩(GVC)框架,在部分情况下达到0.02%的压缩率。该方法重新定义视频压缩极限,利用现代生成视频模型实现极端压缩率,同时保持以感知为中心、任务导向的通信范式,对应香农-韦弗模型中的Level C。GVC通过将计算负担从传输转移到推理:将视频编码为极小表示,由接收端借助强大的生成先验从少量传输信息中合成高质量视频。为确保实际部署,我们提出压缩-计算权衡策略,支持在消费级GPU上快速推理。在AI Flow框架下,GVC为应急救援、远程监控和移动边缘计算等带宽与资源受限环境开辟新可能。实证表明,GVC为构建高效、可扩展、实用的新一代视频通信范式提供了可行路径。
原文摘要 · Abstract (English)
Whether a video can be compressed at an extreme compression rate as low as 0.01%? To this end, we achieve the compression rate as 0.02% at some cases by introducing Generative Video Compression (GVC), a new framework that redefines the limits of video compression by leveraging modern generative video models to achieve extreme compression rates while preserving a perception-centric, task-oriented communication paradigm, corresponding to Level C of the Shannon-Weaver model. Besides, How we trade computation for compression rate or bandwidth? GVC answers this question by shifting the burden from transmission to inference: it encodes video into extremely compact representations and delegates content reconstruction to the receiver, where powerful generative priors synthesize high-quality video from minimal transmitted information. Is GVC practical and deployable? To ensure practical deployment, we propose a compression-computation trade-off strategy, enabling fast inference on consume-grade GPUs. Within the AI Flow framework, GVC opens new possibility for video communication in bandwidth- and resource-constrained environments such as emergency rescue, remote surveillance, and mobile edge computing. Through empirical validation, we demonstrate that GVC offers a viable path toward a new effective, efficient, scalable, and practical video communication paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。