arXiv:2501.12330cs.ITcs.LG2025-01被引 1

揭示深度学习图像压缩与理论极限间的差距及成因

The Gap Between Principle and Practice of Lossy Image Coding

  • 从信息论出发,分析五类导致性能差距的效应
  • 实验证明后三类效应仍存显著改进空间
  • 为下一代图像压缩技术提供理论指引

有损图像编码的理论极限由香农信息论中的率失真函数决定。尽管该函数无法精确刻画,近年来深度学习技术已使实际编码方案逼近此极限。现有学习型编码方案通过联合优化率失真代价,大幅超越传统手写编码。然而,仍存在进一步提升空间。本文识别出理论率失真曲线与当前先进学习编码方案间存在的差距,揭示其由建模、近似、摊销、量化和渐进五种效应共同导致。通过仿真与实验定量评估后三种效应,证实未来有损图像编码技术具有巨大潜力。

原文摘要 · Abstract (English)

Lossy image coding is the art of computing that is principally bounded by the image's rate-distortion function. This bound, though never accurately characterized, has been approached practically via deep learning technologies in recent years. Indeed, learned image coding schemes allow direct optimization of the joint rate-distortion cost, thereby outperforming the handcrafted image coding schemes by a large margin. Still, it is observed that there is room for further improvement in the rate-distortion performance of learned image coding. In this article, we identify the gap between the ideal rate-distortion function forecasted by Shannon's information theory and the empirical rate-distortion function achieved by the state-of-the-art learned image coding schemes, revealing that the gap is incurred by five different effects: modeling effect, approximation effect, amortization effect, digitization effect, and asymptotic effect. We design simulations and experiments to quantitively evaluate the last three effects, which demonstrates the high potential of future lossy image coding technologies.

图像压缩率失真深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。