arXiv:2604.05347eess.IVcs.CV2026-04

为机器视觉优化图像编码,显著提升目标检测与分割性能。

CI-ICM: Channel Importance-driven Learned Image Coding for Machines

  • 按通道重要性动态分配比特,优先保护关键特征
  • 在COCO2017上实现检测与分割任务16.25%和13.72%的性能提升
  • 支持多任务自适应,适合工业级机器视觉应用

传统以人眼为中心的图像压缩方法在机器视觉场景下表现不佳,因视觉特性与特征需求不同。本文提出面向机器视觉的通道重要性驱动学习图像编码(CI-ICM),旨在给定码率约束下最大化机器视觉任务性能。首先,设计通道重要性生成(CIG)模块量化通道重要性,并引入通道排序损失进行降序排列。其次,提出特征通道分组与缩放(FCGS)模块,依据重要性非均匀分组并调整各组动态范围;进一步设计通道重要性上下文(CI-CTX)模块,实现组间比特分配与关键通道高保真保留。第三,提出任务特定通道适配(TSCA)模块,自适应增强多种下游机器任务特征。在COCO2017数据集上的实验表明,相比基准编码器,CI-ICM在目标检测任务中取得16.25%的BD-mAP@50:95增益,在实例分割任务中提升13.72%。消融实验证实各模块有效性,计算复杂度分析表明其具备实际可行性。本工作建立了面向机器视觉的特征通道优化范式,弥合了图像编码与机器感知之间的鸿沟。

原文摘要 · Abstract (English)

Traditional human vision-centric image compression methods are suboptimal for machine vision centric compression due to different visual properties and feature characteristics. To address this problem, we propose a Channel Importance-driven learned Image Coding for Machines (CI-ICM), aiming to maximize the performance of machine vision tasks at a given bitrate constraint. First, we propose a Channel Importance Generation (CIG) module to quantify channel importance in machine vision and develop a channel order loss to rank channels in descending order. Second, to properly allocate bitrate among feature channels, we propose a Feature Channel Grouping and Scaling (FCGS) module that non-uniformly groups the feature channels based on their importance and adjusts the dynamic range of each group. Based on FCGS, we further propose a Channel Importance-based Context (CI-CTX) module to allocate bits among feature groups and to preserve higher fidelity in critical channels. Third, to adapt to multiple machine tasks, we propose a Task-Specific Channel Adaptation (TSCA) module to adaptively enhance features for multiple downstream machine tasks. Experimental results on the COCO2017 dataset show that the proposed CI-ICM achieves BD-mAP@50:95 gains of 16.25$\%$ in object detection and 13.72$\%$ in instance segmentation over the established baseline codec. Ablation studies validate the effectiveness of each contribution, and computation complexity analysis reveals the practicability of the CI-ICM. This work establishes feature channel optimization for machine vision-centric compression, bridging the gap between image coding and machine perception.

图像编码机器视觉通道优化目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。