提出跨模态联合压缩框架,显著降低红外与可见光图像传输开销。
End-to-End RGB-IR Joint Image Compression With Channel-wise Cross-modality Entropy Model
- 设计通道级跨模态熵模型,利用双模态低频信息建模上下文概率。
- 在LLVIP数据集上相比现有方法节省23.1%码率,性能领先。
- 适合需要高效多模态图像传输的智能监控场景使用。
RGB-IR(可见光-红外)图像对广泛应用于智能监控等场景。随着模态数量增加,存储与传输成本翻倍,因此高效压缩至关重要。本文提出一种端到端的RGB-IR图像对联合压缩框架。为充分挖掘跨模态先验信息以实现更精确的上下文概率建模,提出通道级跨模态熵模型(CCEM)。CCEM包含低频上下文提取块(LCEB)和低频上下文融合块(LCFB),用于提取并融合双模态的全局低频信息,辅助更准确地预测熵参数。实验结果表明,该方法在LLVIP和KAIST数据集上均优于现有单模态及多模态压缩方法。例如,在LLVIP数据集上,相比CVPR 2022提出的先进RGB-IR编码器,本框架实现23.1%的码率节省。
原文摘要 · Abstract (English)
RGB-IR(RGB-Infrared) image pairs are frequently applied simultaneously in various applications like intelligent surveillance. However, as the number of modalities increases, the required data storage and transmission costs also double. Therefore, efficient RGB-IR data compression is essential. This work proposes a joint compression framework for RGB-IR image pair. Specifically, to fully utilize cross-modality prior information for accurate context probability modeling within and between modalities, we propose a Channel-wise Cross-modality Entropy Model (CCEM). Among CCEM, a Low-frequency Context Extraction Block (LCEB) and a Low-frequency Context Fusion Block (LCFB) are designed for extracting and aggregating the global low-frequency information from both modalities, which assist the model in predicting entropy parameters more accurately. Experimental results demonstrate that our approach outperforms existing RGB-IR image pair and single-modality compression methods on LLVIP and KAIST datasets. For instance, the proposed framework achieves a 23.1% bit rate saving on LLVIP dataset compared to the state-of-the-art RGB-IR image codec presented at CVPR 2022.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。