开源平台评估视觉任务压缩效果,兼顾画质与模型精度。
CompressAI-Vision: Open-source software to evaluate compression methods for computer vision tasks
- 构建统一评测平台,对比压缩算法在视觉任务中的表现。
- 支持远程与拆分推理场景,量化比特率与任务准确率关系。
- 被MPEG用于机器特征编码标准开发,适合研究者与工业界使用。
随着基于神经网络的计算机视觉应用日益普及,针对视觉任务优化的视频压缩技术受到关注。由于视觉任务、对应模型和数据集多样,亟需一个统一平台来实现并评估面向下游任务的压缩方法。CompressAI-Vision作为综合性评估平台,支持新编码工具在两种推理场景——远程与拆分推理——下,高效压缩视觉网络输入的同时保持任务精度。通过结合标准编码器(开发中),在多个数据集上评估了比特率与任务准确率之间的权衡。该平台为开源软件,已被动态影像专家组(MPEG)采纳,用于推进机器特征编码(FCM)标准的制定。代码公开于 https://github.com/InterDigitalInc/CompressAI-Vision。
原文摘要 · Abstract (English)
With the increasing use of neural network (NN)-based computer vision applications that process image and video data as input, interest has emerged in video compression technology optimized for computer vision tasks. In fact, given the variety of vision tasks, associated NN models and datasets, a consolidated platform is needed as a common ground to implement and evaluate compression methods optimized for downstream vision tasks. CompressAI-Vision is introduced as a comprehensive evaluation platform where new coding tools compete to efficiently compress the input of vision network while retaining task accuracy in the context of two different inference scenarios: "remote" and "split" inferencing. Our study showcases various use cases of the evaluation platform incorporated with standard codecs (under development) by examining the compression gain on several datasets in terms of bit-rate versus task accuracy. This evaluation platform has been developed as open-source software and is adopted by the Moving Pictures Experts Group (MPEG) for the development the Feature Coding for Machines (FCM) standard. The software is available publicly at https://github.com/InterDigitalInc/CompressAI-Vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。