MIRC通过多尺度量化压缩,显著提升图像编码效率。
Multi-scale Image Representation Compression

- 全链路端到端量化,统一优化率失真目标
- 在CLIC2020上比VVC降低10.5%的码率
- 支持1.2~2.9千MAC/像素灵活部署
过拟合编码器在图像与视频压缩中表现优异。以Cool-chic系列模型为例,其在保持较低解码复杂度的同时,性能可媲美通用模型,但需较长过拟合过程。然而,这些编码器未完全优化率失真目标:训练时权重保持全精度,量化参数在后期单独选择,且合成仅在单尺度进行,忽略跨尺度冗余。本文提出MIRC,一种过拟合图像编码器,将潜在表示、合成网络和熵模型均在统一率失真目标下进行量化与熵编码,采用神经视频表示编码器NVRC的端到端流程。进一步引入跨阶段参数共享的多尺度表示,以微小传输开销提升编码效率。在CLIC2020专业验证集上,MIRC相较VVC(VTM 22.0)实现10.5%的BD-rate节省。同时,MIRC提供1.2至2.9 kMAC/像素的多种配置,可根据部署需求灵活选择解码预算。
原文摘要 · Abstract (English)
Overfitted codecs have demonstrated promising performance for image and video compression. In particular, for image compression, the Cool-chic family of models has shown competitive performance against scene-agnostic models, with orders of magnitude lower decoding complexity at the cost of a longer overfitting process. However, these overfitted image codecs are not fully optimized toward the rate-distortion objective: their network weights remain in full precision during training, and the associated quantization parameters are selected in a separate post-training stage. Furthermore, their synthesis operates at a single scale, which overlooks cross-scale redundancy. In this paper, we propose MIRC, an overfitted image codec in which every coded component, including the latents, the synthesis network, and the entropy models, is quantized and entropy coded under a single rate-distortion objective, adopting the end-to-end compression pipeline of the neural video representation codec NVRC. We further introduce a multi-scale representation with cross-stage parameter sharing, which improves coding efficiency at a small transmitted overhead. On the CLIC2020 professional validation set, MIRC achieves a 10.5% BD-rate saving against VVC (VTM 22.0). Moreover, MIRC offers a family of configurations spanning 1.2 to 2.9 kMAC per pixel, so the decoding budget can be selected to match the deployment target.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。