用感知优化替代数据建模,实现高质量低算力图像压缩。
Good, Cheap, and Fast: Overfitted Image Compression with Wasserstein Distortion
- 以感知误差为优化目标,而非建模图像分布
- 压缩耗时低于商用编码器1%,质量接近生成式模型
- 感知指标与人评相关性超94%,优于主流度量
受生成式图像模型成功启发,当前学习型图像压缩多聚焦于自然图像分布的概率建模,虽能获得优异视觉质量,但计算复杂度比现有商用编码器高数个数量级,难以实用。本文提出通过聚焦视觉感知建模,而非数据分布建模,使过拟合的C3编码器在优化Wasserstein Distortion(WD)时,实现与HiFiC等生成式压缩模型相当的画质与码率平衡,同时解码所需乘加操作(MACs)不足1%。通过人工评分实验验证,WD作为优化目标显著优于LPIPS;且在预测人类主观评价方面,其与Elo评分的皮尔逊相关系数超过94%,优于LPIPS、DISTS和MS-SSIM等主流感知度量。
原文摘要 · Abstract (English)
Inspired by the success of generative image models, recent work on learned image compression increasingly focuses on better probabilistic models of the natural image distribution, leading to excellent image quality. This, however, comes at the expense of a computational complexity that is several orders of magnitude higher than today's commercial codecs, and thus prohibitive for most practical applications. With this paper, we demonstrate that by focusing on modeling visual perception rather than the data distribution, we can achieve a very good trade-off between visual quality and bit rate similar to "generative" compression models such as HiFiC, while requiring less than 1% of the multiply-accumulate operations (MACs) for decompression. We do this by optimizing C3, an overfitted image codec, for Wasserstein Distortion (WD), and evaluating the image reconstructions with a human rater study, showing that WD clearly outperforms LPIPS as an optimization objective. The study also reveals that WD outperforms other perceptual metrics such as LPIPS, DISTS, and MS-SSIM as a predictor of human ratings, remarkably achieving over 94% Pearson correlation with Elo scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。