为机器设计的立体图像压缩网络,提升3D视觉任务效率
Stereo Image Coding for Machines with Joint Visual Feature Compression
- 构建面向机器视觉的立体特征压缩网络,联合优化多尺度视图特征
- 在相同码率下,3D任务性能优于现有标准和先进方法
- 适合自动驾驶、机器人等需要高效立体视觉分析的场景
针对机器的二维图像编码(ICM)已在编码效率上取得显著进展,但立体图像领域研究较少。本文提出机器视觉导向的立体图像编码(SICM),并设计了多视图立体特征压缩网络(MVSFC-Net),用于高效提取、压缩与传输立体视觉特征以支持三维视觉任务。为此,提出一种立体多尺度特征压缩(SMFC)模块,通过同时消除空间、跨视图及跨尺度冗余,将稀疏的多尺度立体特征逐步转换为紧凑的联合视觉表示。实验表明,所提MVSFC-Net在压缩效率和3D视觉任务性能方面均优于MPEG推荐的ICM基准及当前最优的立体图像压缩方法。
原文摘要 · Abstract (English)
2D image coding for machines (ICM) has achieved great success in coding efficiency, while less effort has been devoted to stereo image fields. To promote the efficiency of stereo image compression (SIC) and intelligent analysis, the stereo image coding for machines (SICM) is formulated and explored in this paper. More specifically, a machine vision-oriented stereo feature compression network (MVSFC-Net) is proposed for SICM, where the stereo visual features are effectively extracted, compressed, and transmitted for 3D visual task. To efficiently compress stereo visual features in MVSFC-Net, a stereo multi-scale feature compression (SMFC) module is designed to gradually transform sparse stereo multi-scale features into compact joint visual representations by removing spatial, inter-view, and cross-scale redundancies simultaneously. Experimental results show that the proposed MVSFC-Net obtains superior compression efficiency as well as 3D visual task performance, when compared with the existing ICM anchors recommended by MPEG and the state-of-the-art SIC method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。