arXiv:2412.03841cs.CVcs.AI2024-12

用大模型提升低层机器视觉图像压缩,让编码与任务协同优化。

LL-ICM: Image Compression for Low-level Machine Vision via Large Vision-Language Model

  • 联合优化压缩与低层视觉任务,实现编码与模型互适。
  • 在多个任务上实现22.65%的码率降低,优于现有方法。
  • 引入视觉语言模型生成鲁棒特征,支持多任务泛化。

图像压缩面向机器(ICM)旨在为机器视觉任务压缩图像而非人眼观看。当前研究多聚焦于高层任务如目标检测和语义分割,但真实场景中原始图像质量常不理想,压缩后性能进一步下降。低层(LL)机器视觉模型如图像恢复可改善质量,因此其压缩需求也应被考虑。本文提出首个面向低层机器视觉任务的ICM框架——LL-ICM。通过联合优化压缩与低层任务,该框架不仅增强编码对多样化低层任务的泛化能力,还优化下游低层任务模型的处理性能,实现编码器与任务模型的协同适应。此外,将大规模视觉语言模型融入框架,生成更通用且抗失真的特征嵌入,使单一LL-ICM编码器可泛化至多种任务。我们建立基准评测体系,包含全参考与无参考图像质量评估。实验表明,LL-ICM相比最先进方法实现22.65%的BD-rate降低。

原文摘要 · Abstract (English)

Image Compression for Machines (ICM) aims to compress images for machine vision tasks rather than human viewing. Current works predominantly concentrate on high-level tasks like object detection and semantic segmentation. However, the quality of original images is usually not guaranteed in the real world, leading to even worse perceptual quality or downstream task performance after compression. Low-level (LL) machine vision models, like image restoration models, can help improve such quality, and thereby their compression requirements should also be considered. In this paper, we propose a pioneered ICM framework for LL machine vision tasks, namely LL-ICM. By jointly optimizing compression and LL tasks, the proposed LL-ICM not only enriches its encoding ability in generalizing to versatile LL tasks but also optimizes the processing ability of down-stream LL task models, achieving mutual adaptation for image codecs and LL task models. Furthermore, we integrate large-scale vision-language models into the LL-ICM framework to generate more universal and distortion-robust feature embeddings for LL vision tasks. Therefore, one LL-ICM codec can generalize to multiple tasks. We establish a solid benchmark to evaluate LL-ICM, which includes extensive objective experiments by using both full and no-reference image quality assessments. Experimental results show that LL-ICM can achieve 22.65% BD-rate reductions over the state-of-the-art methods.

图像压缩低层视觉视觉语言模型机器感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。