arXiv:2505.17366eess.IV2025-05

用低秩微调让预训练模型高效适配多任务图像编码,省电又省空间。

Low-Rank Adaptation of Pre-trained Vision Backbones for Energy-Efficient Image Coding for Machine

  • 用低秩适配技术微调预训练视觉模型,保留主干特征同时降低参数量。
  • 在多个数据集上实现比传统编解码器更高的压缩效率,且不需全量微调。
  • 适合需要多任务、低功耗的机器视觉应用,如自动驾驶与边缘计算。

机器图像编码(ICM)旨在优化图像压缩以服务人工智能分析,而非人眼感知。现有ICM框架通常为特定任务设计独立编解码器,导致存储开销大、训练成本高、计算复杂度高。为此,本文提出一种节能型框架,利用预训练视觉主干提取适用于多任务的鲁棒、通用潜在表示。引入任务特定的低秩适配机制,对预训练特征进行精细化调整,使其既可压缩又适配下游应用。该设计大幅减少可训练参数,降低多任务场景下的能耗。通过联合优化任务性能与熵最小化,方法可在无需全量微调的情况下高效适配多种任务与数据集,实现高编码效率。大量实验表明,本框架显著优于传统编解码器与预处理器,为ICM应用提供了一种节能高效的解决方案。代码与补充材料将公开于:https://gitlab.com/viper-purdue/efficient-compression。

原文摘要 · Abstract (English)

Image Coding for Machines (ICM) focuses on optimizing image compression for AI-driven analysis rather than human perception. Existing ICM frameworks often rely on separate codecs for specific tasks, leading to significant storage requirements, training overhead, and computational complexity. To address these challenges, we propose an energy-efficient framework that leverages pre-trained vision backbones to extract robust and versatile latent representations suitable for multiple tasks. We introduce a task-specific low-rank adaptation mechanism, which refines the pre-trained features to be both compressible and tailored to downstream applications. This design minimizes trainable parameters and reduces energy costs for multi-task scenarios. By jointly optimizing task performance and entropy minimization, our method enables efficient adaptation to diverse tasks and datasets without full fine-tuning, achieving high coding efficiency. Extensive experiments demonstrate that our framework significantly outperforms traditional codecs and pre-processors, offering an energy-efficient and effective solution for ICM applications. The code and the supplementary materials will be available at: https://gitlab.com/viper-purdue/efficient-compression.

图像编码低秩适配预训练模型节能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。