arXiv:2512.15270eess.IVcs.CV2025-12中稿 · as a PAPER and for…

用预训练扩散模型做图像压缩前处理,提升画质同时大幅降低码率。

Generative Preprocessing for Image Compression with Pre-trained Diffusion Models

  • 用一致性分数蒸馏将扩散模型压缩为单步图像生成器。
  • 在Kodak数据集上实现30.13%的BD-rate降低,主观画质更优。
  • 无需修改现有编码器,适合追求高质量低码率的图像压缩场景。

预处理是优化图像压缩的有效手段,但现有方法主要基于率失真(R-D)准则,受限于像素级保真度。本文首次提出转向率感知(R-P)优化,利用大规模预训练扩散模型进行压缩前处理。提出两阶段框架:首先通过一致分数身份蒸馏(CiD)将多步Stable Diffusion 2.1压缩为轻量单步图像到图像模型;其次,在可参数高效微调的注意力模块上,以率感知损失和可微编码器代理为指导进行优化。该方法无缝集成于标准编码器,无需修改,利用模型强大的生成先验增强纹理并抑制伪影。实验显示显著的R-P性能提升,在Kodak数据集上DISTS指标下实现最高30.13%的BD-rate降低,并获得更优主观视觉质量。

原文摘要 · Abstract (English)

Preprocessing is a well-established technique for optimizing compression, yet existing methods are predominantly Rate-Distortion (R-D) optimized and constrained by pixel-level fidelity. This work pioneers a shift towards Rate-Perception (R-P) optimization by, for the first time, adapting a large-scale pre-trained diffusion model for compression preprocessing. We propose a two-stage framework: first, we distill the multi-step Stable Diffusion 2.1 into a compact, one-step image-to-image model using Consistent Score Identity Distillation (CiD). Second, we perform a parameter-efficient fine-tuning of the distilled model's attention modules, guided by a Rate-Perception loss and a differentiable codec surrogate. Our method seamlessly integrates with standard codecs without any modification and leverages the model's powerful generative priors to enhance texture and mitigate artifacts. Experiments show substantial R-P gains, achieving up to a 30.13% BD-rate reduction in DISTS on the Kodak dataset and delivering superior subjective visual quality.

图像压缩扩散模型率感知生成式预处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。