用扩散模型先验提升低比特率图像压缩的感知质量。
Generative Image Coding with Diffusion Prior
- 用预优化编码器结合扩散模型特征,轻量适配实现高效压缩。
- 在低码率下重建质量显著优于现有方法,最高压缩效率提升79%。
- 适合处理AI生成内容,且可低成本适配不同模型需求。
随着生成技术发展,视觉内容日益混合自然与AI生成图像,亟需更高效的编码方法以保障感知质量。传统编解码器和学习方法在高压缩比下难以维持主观质量,现有生成式方法则面临视觉保真度与泛化能力不足的问题。为此,我们提出一种基于扩散先验的生成式编码框架,利用预优化编码器生成通用压缩域表示,通过轻量适配器与注意力融合模块整合预训练扩散模型内部特征。该框架有效利用现有预训练模型,支持低重训成本适配新需求。此外,引入分布重校准方法进一步提升重建保真度。大量实验表明:(1) 在低比特率下视觉保真度全面领先;(2) 压缩性能相比H.266/VVC最高提升79%;(3) 为AI生成内容提供高效解决方案,并兼容多种内容类型。
原文摘要 · Abstract (English)
As generative technologies advance, visual content has evolved into a complex mix of natural and AI-generated images, driving the need for more efficient coding techniques that prioritize perceptual quality. Traditional codecs and learned methods struggle to maintain subjective quality at high compression ratios, while existing generative approaches face challenges in visual fidelity and generalization. To this end, we propose a novel generative coding framework leveraging diffusion priors to enhance compression performance at low bitrates. Our approach employs a pre-optimized encoder to generate generalized compressed-domain representations, integrated with the pretrained model's internal features via a lightweight adapter and an attentive fusion module. This framework effectively leverages existing pretrained diffusion models and enables efficient adaptation to different pretrained models for new requirements with minimal retraining costs. We also introduce a distribution renormalization method to further enhance reconstruction fidelity. Extensive experiments show that our method (1) outperforms existing methods in visual fidelity across low bitrates, (2) improves compression performance by up to 79% over H.266/VVC, and (3) offers an efficient solution for AI-generated content while being adaptable to broader content types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。