arXiv:2412.15752cs.CVeess.IV2024-12中稿 · TCSVT被引 2

用点云辅助图像压缩,提升自动驾驶场景下的编码效率。

Sparse Point Clouds Assisted Learned Image Compression

  • 将3D点云投影为稀疏深度图,用于预测图像并提取多尺度结构特征。
  • 在多个主流压缩模型上验证,均实现稳定性能提升。
  • 适合自动驾驶中多模态数据融合的图像压缩任务。

在自动驾驶领域,多种传感器数据代表同一场景的不同模态,因此可利用其他传感器数据辅助图像压缩。然而,目前很少有技术探索利用跨模态相关性来提升图像压缩性能。本文受近年来学习型图像压缩成功启发,提出一种新框架,利用稀疏点云辅助自动驾驶场景下的学习型图像压缩。首先将3D稀疏点云投影至2D平面,生成稀疏深度图;基于该深度图预测相机图像,并提取多尺度结构特征;这些特征作为额外信息融入学习型图像压缩流程,以提升压缩性能。所提框架兼容多种主流学习型图像压缩模型,实验验证了其在不同压缩方法上的有效性。结果表明,引入点云辅助可持续提升压缩性能。

原文摘要 · Abstract (English)

In the field of autonomous driving, a variety of sensor data types exist, each representing different modalities of the same scene. Therefore, it is feasible to utilize data from other sensors to facilitate image compression. However, few techniques have explored the potential benefits of utilizing inter-modality correlations to enhance the image compression performance. In this paper, motivated by the recent success of learned image compression, we propose a new framework that uses sparse point clouds to assist in learned image compression in the autonomous driving scenario. We first project the 3D sparse point cloud onto a 2D plane, resulting in a sparse depth map. Utilizing this depth map, we proceed to predict camera images. Subsequently, we use these predicted images to extract multi-scale structural features. These features are then incorporated into learned image compression pipeline as additional information to improve the compression performance. Our proposed framework is compatible with various mainstream learned image compression models, and we validate our approach using different existing image compression methods. The experimental results show that incorporating point cloud assistance into the compression pipeline consistently enhances the performance.

图像压缩点云自动驾驶多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。