为机器视觉优化视频压缩,降低15%码率仍保持精度
A Preprocessing Framework for Video Machine Vision under Compression
- 用神经预处理器保留关键信息,提升压缩效率
- 在多种模型上实现超15%码率节省,精度不降
- 兼容标准编码器,适合真实场景部署
随着终端视频压缩与传输在机器视觉任务中日益普遍,现有视频编码优化多基于人眼感知指标最小化失真,忽视了机器视觉系统的更高需求。本文提出一种面向机器视觉的视频预处理框架,通过引入神经预处理器保留下游任务关键信息,显著提升率-精度性能。我们进一步设计可微分虚拟编码器,在训练阶段施加码率与失真约束,并直接使用主流标准编码器进行测试。实验在两类典型下游任务及多种骨干网络上验证,结果表明,相比仅使用标准编码器的基准版本,本方法可节省超过15%码率。
原文摘要 · Abstract (English)
There has been a growing trend in compressing and transmitting videos from terminals for machine vision tasks. Nevertheless, most video coding optimization method focus on minimizing distortion according to human perceptual metrics, overlooking the heightened demands posed by machine vision systems. In this paper, we propose a video preprocessing framework tailored for machine vision tasks to address this challenge. The proposed method incorporates a neural preprocessor which retaining crucial information for subsequent tasks, resulting in the boosting of rate-accuracy performance. We further introduce a differentiable virtual codec to provide constraints on rate and distortion during the training stage. We directly apply widely used standard codecs for testing. Therefore, our solution can be easily applied to real-world scenarios. We conducted extensive experiments evaluating our compression method on two typical downstream tasks with various backbone networks. The experimental results indicate that our approach can save over 15% of bitrate compared to using only the standard codec anchor version.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。