提出兼顾视觉质量与运行速度的实用化图像压缩新方法。
What Matters in Practical Learned Image Compression

- 联合优化感知质量与运行效率,设计可实用的神经编码器。
- 在1200万像素图像上编码仅需230毫秒,解码150毫秒。
- 相比AV1等主流标准,压缩率提升2.3至3倍,适合移动端部署。
相较于传统硬编码方案,学习型编码器的一大优势在于能直接针对人类视觉系统进行优化。然而,至今尚未有兼顾感知质量与实用性的图像编码方案被提出。本文旨在填补这一空白,系统研究影响实际学习型图像编码器设计的关键建模选择,包括多项新提出的优化技术,并联合优化感知质量与运行时性能。通过在数百万种骨干网络配置中进行性能感知的神经架构搜索,识别出在目标设备端运行时延下实现最佳压缩性能的模型。最终构建的新编码器显著提升了速度与感知质量之间的权衡。基于严格的主观用户测试,该方法在比特率上相较AV1、AV2、VVC、ECM和JPEG-AI降低了2.3-3倍,比最优学习型编码器节省20%-40%比特率。在iPhone 17 Pro Max上,1200万像素图像的编码时间仅为230毫秒,解码时间为150毫秒,快于大多数基于ML的编码器在V100 GPU上的表现。
原文摘要 · Abstract (English)
One of the major differentiators unlocked by learned codecs relative to their hard-coded traditional counterparts is their ability to be optimized directly to appeal to the human visual system. Despite this potential, a perceptual yet practical image codec is yet to be proposed. In this work, we aim to close this gap. We conduct a comprehensive study of the key modeling choices that govern the design of a practical learned image codec, jointly optimized for perceptual quality and runtime -- including within the ablations several novel techniques. We then perform performance-aware neural architecture search over millions of backbone configurations to identify models that achieve the target on-device runtime while maximizing compression performance as captured by perceptual metrics. We combine the various optimizations to construct a new codec that achieves a significantly improved tradeoff between speed and perceptual quality. Based on rigorous subjective user studies, it provides 2.3-3x bitrate savings against AV1, AV2, VVC, ECM and JPEG-AI, and 20-40% bitrate savings against the best learned codec alternatives. At the same time, on an iPhone 17 Pro Max, it encodes 12MP images as fast as 230ms, and decodes them in 150ms -- faster than most top ML-based codecs run on a V100 GPU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。