为机器视觉任务优化视频编码,提升特征压缩效率与速度。
New VVC profiles targeting Feature Coding for Machines
- 针对机器任务设计轻量VVC编码方案,聚焦中间特征压缩。
- 最快版本编码时间减少95.6%,仅损失1.71%压缩率。
- 适用于边缘计算、模型分发等需要高效传输特征的场景。
现代视频编码器主要优化人眼感知质量,依赖人类视觉系统模型。但在分割推理系统中,传输的是神经网络的中间特征而非像素数据,这些特征抽象、稀疏且任务特定,感知保真度不再相关。本文研究在MPEG-AI机器特征编码(FCM)标准下,使用通用视频编码(VVC)压缩此类特征的可行性。通过工具级分析,评估各编码组件对压缩效率和下游视觉任务准确率的影响。基于此,提出三个轻量级核心VVC配置:Fast、Faster和Fastest。Fast配置实现2.96%的BD-Rate提升,编码时间减少21.8%;Faster在保持1.85% BD-Rate增益的同时,编码速度提升51.5%;Fastest将编码时间降低95.6%,仅导致1.71%的BD-Rate性能下降。
原文摘要 · Abstract (English)
Modern video codecs have been extensively optimized to preserve perceptual quality, leveraging models of the human visual system. However, in split inference systems-where intermediate features from neural network are transmitted instead of pixel data-these assumptions no longer apply. Intermediate features are abstract, sparse, and task-specific, making perceptual fidelity irrelevant. In this paper, we investigate the use of Versatile Video Coding (VVC) for compressing such features under the MPEG-AI Feature Coding for Machines (FCM) standard. We perform a tool-level analysis to understand the impact of individual coding components on compression efficiency and downstream vision task accuracy. Based on these insights, we propose three lightweight essential VVC profiles-Fast, Faster, and Fastest. The Fast profile provides 2.96% BD-Rate gain while reducing encoding time by 21.8%. Faster achieves a 1.85% BD-Rate gain with a 51.5% speedup. Fastest reduces encoding time by 95.6% with only a 1.71% loss in BD-Rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。