用扩散模型统一压缩屏幕与自然图像,文本更清晰。
PICD: Versatile Perceptual Image Compression with Diffusion Rendering
- 分离编码图文,用扩散模型融合渲染
- 在三个层次注入条件信息,提升文本还原度
- 兼顾屏幕与自然图像,适合多场景应用
近年来,感知图像压缩在自然图像上取得了显著进展,可在低码率下实现高视觉质量。然而,现有方法在压缩屏幕内容时,尤其是文字部分,常产生明显伪影。为此,我们提出兼具通用性的感知屏幕图像压缩框架(PICD),该框架能有效处理屏幕与自然图像。具体而言,我们设计了一种分别编码文本与图像的压缩架构,并利用扩散模型将二者融合为一张图像。在扩散渲染中,我们在三个层面集成条件信息:1)领域级:使用屏幕内容提示微调基础扩散模型;2)适配器级:开发高效适配器,以压缩后的图像和文本作为输入控制扩散过程;3)实例级:应用实例级引导进一步优化解码。实验表明,我们的PICD在文本准确性和感知质量方面均优于现有感知编码器。此外,在无文本条件时,该方法仍可作为自然图像的高效感知编码器。
原文摘要 · Abstract (English)
Recently, perceptual image compression has achieved significant advancements, delivering high visual quality at low bitrates for natural images. However, for screen content, existing methods often produce noticeable artifacts when compressing text. To tackle this challenge, we propose versatile perceptual screen image compression with diffusion rendering (PICD), a codec that works well for both screen and natural images. More specifically, we propose a compression framework that encodes the text and image separately, and renders them into one image using diffusion model. For this diffusion rendering, we integrate conditional information into diffusion models at three distinct levels: 1). Domain level: We fine-tune the base diffusion model using text content prompts with screen content. 2). Adaptor level: We develop an efficient adaptor to control the diffusion model using compressed image and text as input. 3). Instance level: We apply instance-wise guidance to further enhance the decoding process. Empirically, our PICD surpasses existing perceptual codecs in terms of both text accuracy and perceptual quality. Additionally, without text conditions, our approach serves effectively as a perceptual codec for natural images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。