提出分层语义压缩框架,实现高效且一致的图像语义恢复。
Hierarchical Semantic Compression for Consistent Image Semantic Restoration
- 基于生成模型内在语义空间,分层压缩中层特征与核心语义。
- 在极低码率下保持人类视觉主观质量和语义一致性领先。
- 适合关注图像压缩与语义保持的研究者和应用开发者。
新兴的语义压缩近年来受到广泛关注,能在极低码率下实现高保真恢复。然而,现有方法通常结合预定义或高维语义的标准流程,压缩效率受限。为此,我们提出一种纯在生成模型内在语义空间操作的分层语义压缩(HSC)框架,实现高效且一致的语义恢复。具体地,我们分析了语义压缩的熵模型,据此设计基于新开发通用反演编码器的分层架构;提出特征压缩网络(FCN)与语义压缩网络(SCN),通过逐通道上下文共享的熵模型,分层压缩中间语义特征与核心语义,以恢复语义准确性和一致性。实验表明,所提HSC框架在主观质量与语义一致性上达到当前最优表现,且在机器视觉任务中也展现出优越性能,与人类视觉系统理解图像的方式高度契合,为未来图像/视频压缩范式提供新思路。代码将在接受后发布。
原文摘要 · Abstract (English)
The emerging semantic compression has been receiving increasing research efforts most recently, capable of achieving high fidelity restoration during compression, even at extremely low bitrates. However, existing semantic compression methods typically combine standard pipelines with either pre-defined or high-dimensional semantics, thus suffering from deficiency in compression. To address this issue, we propose a novel hierarchical semantic compression (HSC) framework that purely operates within intrinsic semantic spaces from generative models, which is able to achieve efficient compression for consistent semantic restoration. More specifically, we first analyse the entropy models for the semantic compression, which motivates us to employ a hierarchical architecture based on a newly developed general inversion encoder. Then, we propose the feature compression network (FCN) and semantic compression network (SCN), such that the middle-level semantic feature and core semantics are hierarchically compressed to restore both accuracy and consistency of image semantics, via an entropy model progressively shared by channel-wise context. Experimental results demonstrate that the proposed HSC framework achieves the state-of-the-art performance on subjective quality and consistency for human vision, together with superior performances on machine vision tasks given compressed bitstreams. This essentially coincides with human visual system in understanding images, thus providing a new framework for future image/video compression paradigms. Our code shall be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。