根据特征尺度重要性分配比特,显著降低机器视觉特征编码的带宽需求。
Multiscale Feature Importance-based Bit Allocation for End-to-End Feature Coding for Machines
- 基于多尺度特征重要性动态分配编码比特,提升压缩效率。
- 在目标检测上平均节省38.2%码率,实例分割和关键点检测分别节省17.2%和36.5%。
- 适用于多种视觉任务与编码器,通用性强,适合远程智能分析场景。
机器特征编码(FCM)旨在高效压缩中间特征以支持远程智能分析,对未来智能视觉应用至关重要。本文提出一种多尺度特征重要性比特分配方法(MFIBA),用于端到端的FCM。首先,发现机器视觉任务中特征的重要性随尺度、物体大小和图像实例变化,据此提出多尺度特征重要性预测(MFIP)模块,为各尺度特征预测重要性权重。其次,构建任务损失-码率模型,建立使用压缩特征时任务精度损失与编码码率之间的关系。最后,设计基于重要性的端到端比特分配策略,实现更合理的多尺度特征编码。实验表明,结合保留的ELIC编码器,所提方法在目标检测中平均节省38.202%码率;在实例分割和关键点检测中分别节省17.212%和36.492%。应用于LIC-TCM时,三种任务平均码率节省分别为18.103%、19.866%和19.597%,验证了其在不同任务和编码器间的良好泛化性与适应性。
原文摘要 · Abstract (English)
Feature Coding for Machines (FCM) aims to compress intermediate features effectively for remote intelligent analytics, which is crucial for future intelligent visual applications. In this paper, we propose a Multiscale Feature Importance-based Bit Allocation (MFIBA) for end-to-end FCM. First, we find that the importance of features for machine vision tasks varies with the scales, object size, and image instances. Based on this finding, we propose a Multiscale Feature Importance Prediction (MFIP) module to predict the importance weight for each scale of features. Secondly, we propose a task loss-rate model to establish the relationship between the task accuracy losses of using compressed features and the bitrate of encoding these features. Finally, we develop a MFIBA for end-to-end FCM, which is able to assign coding bits of multiscale features more reasonably based on their importance. Experimental results demonstrate that when combined with a retained Efficient Learned Image Compression (ELIC), the proposed MFIBA achieves an average of 38.202% bitrate savings in object detection compared to the anchor ELIC. Moreover, the proposed MFIBA achieves an average of 17.212% and 36.492% feature bitrate savings for instance segmentation and keypoint detection, respectively. When the proposed MFIBA is applied to the LIC-TCM, it achieves an average of 18.103%, 19.866% and 19.597% bit rate savings on three machine vision tasks, respectively, which validates the proposed MFIBA has good generalizability and adaptability to different machine vision tasks and FCM base codecs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。