提出新型残差编码机制,提升人机共用图像压缩效率
Explicit Residual-Based Scalable Image Coding for Humans and Machines
- 引入显式残差压缩,借鉴JPEG2000思路提升可解释性
- PR-ICMH实现最高29.57%的BD-rate节省,性能领先
- 兼顾编码复杂度与压缩率,适配多种视觉任务需求
可扩展图像压缩能逐步重建满足不同需求的多版本图像。近年来,图像不仅被人类消费,也被图像识别模型使用,推动了服务人机视觉(ICMH)的可扩展压缩方法发展。现有方法多采用神经网络编码器,即学习型图像压缩,通过精心设计损失函数取得进展,但部分模型过度依赖学习能力,架构设计不足。本文通过集成显式残差压缩机制,提升ICMH框架的编码效率与可解释性,该机制常见于如JPEG2000等分辨率可扩展编码方法。具体提出两种互补方法:基于特征残差的可扩展编码(FR-ICMH)和基于像素残差的可扩展编码(PR-ICMH),适用于多种机器视觉任务。二者在编码复杂度与压缩性能间提供灵活权衡,适应多样化应用需求。实验表明,所提方法有效,其中PR-ICMH相较先前工作最多实现29.57%的BD-rate节省。
原文摘要 · Abstract (English)
Scalable image compression is a technique that progressively reconstructs multiple versions of an image for different requirements. In recent years, images have increasingly been consumed not only by humans but also by image recognition models. This shift has drawn growing attention to scalable image compression methods that serve both machine and human vision (ICMH). Many existing models employ neural network-based codecs, known as learned image compression, and have made significant strides in this field by carefully designing the loss functions. In some cases, however, models are overly reliant on their learning capacity, and their architectural design is not sufficiently considered. In this paper, we enhance the coding efficiency and interpretability of ICMH framework by integrating an explicit residual compression mechanism, which is commonly employed in resolution scalable coding methods such as JPEG2000. Specifically, we propose two complementary methods: Feature Residual-based Scalable Coding (FR-ICMH) and Pixel Residual-based Scalable Coding (PR-ICMH). These proposed methods are applicable to various machine vision tasks. Moreover, they provide flexibility to choose between encoder complexity and compression performance, making it adaptable to diverse application requirements. Experimental results demonstrate the effectiveness of our proposed methods, with PR-ICMH achieving up to 29.57% BD-rate savings over the previous work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。