arXiv:2512.11134eess.IV2025-12被引 1

通过动态保留关键通道,提升机器视觉特征压缩效率。

Feature Compression for Machines with Range-Based Channel Truncation and Frame Packing

  • 根据推理时特征统计动态截断并打包通道。
  • 在保持任务性能前提下平均降低10.59%码率。
  • 适合分布式视觉模型中的特征传输场景。

本文提出一种增强方法,用于即将发布的MPEG特征编码标准(FCM)中机器特征的压缩。该标准旨在实现分计算场景下的互操作压缩比特流,即大型计算机视觉神经网络模型在两个设备间协同推理时,对中间特征进行压缩传输。中间特征通常为多个3D张量,可通过降维与熵编码减少带宽需求。在预设的MPEG-FCM标准设计中,中间特征张量先经神经层压缩,再转化为2D视频帧,利用现有视频压缩标准进行编码。本文引入一种基于范围的通道截断与帧打包方法,使系统能在推理时根据特征统计保留相关通道,同时维持接收端的任务性能。在MPEG-FCM测试模型中实现,该方法在多个视觉任务和数据集上,在保持给定精度的前提下,平均码率降低10.59%。

原文摘要 · Abstract (English)

This paper proposes a method that enhances the compression performance of the current model under development for the upcoming MPEG standard on Feature Coding for Machines (FCM). This standard aims at providing inter-operable compressed bitstreams of features in the context of split computing, i.e., when the inference of a large computer vision neural-network (NN)-based model is split between two devices. Intermediate features can consist of multiple 3D tensors that can be reduced and entropy coded to limit the required bandwidth of such transmission. In the envisioned design for the MPEG-FCM standard, intermediate feature tensors may be reduced using Neural layers before being converted into 2D video frames that can be coded using existing video compression standards. This paper introduces an additional channel truncation and packing method which enables the system to preserve the relevant channels, depending on the statistics of the features at inference time, while preserving the computer vision task performance at the receiver. Implemented within the MPEG-FCM test model, the proposed method yields an average reduction in rate by 10.59% for a given accuracy on multiple computer vision tasks and datasets.

特征压缩分计算视频编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。