用N元语法上下文增强Swin Transformer,实现单模型可变率图像压缩。
Variable Rate Image Compression via N-Gram Context based Swin-transformer
- 引入N元语法上下文提升Swin Transformer的局部感知能力
- 高分辨率重建时BD-Rate降低5.86%,质量显著提升
- 适合工业质检等关注重点区域的应用场景
本文提出一种基于N-gram上下文的Swin Transformer,用于学习型图像压缩。该方法仅用一个模型即可实现可变率压缩。通过将N-gram上下文融入Swin Transformer,克服了其因感受野受限而忽略大区域的问题,扩展了像素恢复的考虑范围,从而提升高分辨率重构质量。该方法增强了相邻窗口间的上下文感知能力,在可变率学习型图像压缩中实现-5.86%的BD-Rate改善。此外,模型在图像感兴趣区域(ROI)上也表现出更优的重建质量,特别适用于制造与工业视觉系统中的目标聚焦应用。
原文摘要 · Abstract (English)
This paper presents an N-gram context-based Swin Transformer for learned image compression. Our method achieves variable-rate compression with a single model. By incorporating N-gram context into the Swin Transformer, we overcome its limitation of neglecting larger regions during high-resolution image reconstruction due to its restricted receptive field. This enhancement expands the regions considered for pixel restoration, thereby improving the quality of high-resolution reconstructions. Our method increases context awareness across neighboring windows, leading to a -5.86\% improvement in BD-Rate over existing variable-rate learned image compression techniques. Additionally, our model improves the quality of regions of interest (ROI) in images, making it particularly beneficial for object-focused applications in fields such as manufacturing and industrial vision systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。