用大模型生成图像描述并编码,实现超低码率下高质量图像压缩。
LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression
- 用大模型统一生成图像描述并压缩,端到端完成语义编码。
- 在相同码率下,相比现有方法提升41.58%的感知质量(LPIPS BD-rate)。
- 适用于追求极致压缩比与视觉质量平衡的研究者或应用开发者。
得益于强大生成模型的支持,基于感知指标的低码率学习型图像压缩(LIC)模型已成为可能。一些先进模型通过利用图像字幕作为辅助信息,实现了高压缩率和优越的感知质量。本文证明,借助大型多模态模型(LMM),可在单一模型中生成字幕并将其压缩。我们还提出一种新型面向语义-感知的微调方法,适用于任何LIC网络,在LPIPS BD-rate上相较现有方法提升41.58%。代码与预训练权重已公开于 https://github.com/tokkiwa/ImageTextCoding。
原文摘要 · Abstract (English)
Supported by powerful generative models, low-bitrate learned image compression (LIC) models utilizing perceptual metrics have become feasible. Some of the most advanced models achieve high compression rates and superior perceptual quality by using image captions as sub-information. This paper demonstrates that using a large multi-modal model (LMM), it is possible to generate captions and compress them within a single model. We also propose a novel semantic-perceptual-oriented fine-tuning method applicable to any LIC network, resulting in a 41.58\% improvement in LPIPS BD-rate compared to existing methods. Our implementation and pre-trained weights are available at https://github.com/tokkiwa/ImageTextCoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。