针对语义特征压缩,提出自适应变换编码方法。
Adaptive Transform Coding for Semantic Compression
- 根据特征分布选择不同变换与量化器,动态适配异质数据
- 在主流视觉模型特征上实现更优压缩率,性能超越或持平顶尖方法
- 兼顾效率与可解释性,适合需要透明压缩的场景
视觉数据压缩正从以人为中心的重建转向面向机器的表征编码。在此背景下,图像常被映射为紧凑的语义嵌入,再进行压缩和传输以支持下游推理。本文提出一种基于高斯混合模型条件率失真函数的自适应变换编码方法,通过依据推断出的源分量选择模式相关的变换与量化器,实现对异质特征分布的更高效编码。在广泛使用的视觉主干网络及基础模型的特征上进行评估,结果表明该方法在压缩性能上优于或媲美当前最先进的神经压缩方法,同时保持灵活性与可解释性。
原文摘要 · Abstract (English)
Visual data compression is shifting from human-centered reconstruction to machine-oriented representation coding. In this setting, an image is often mapped to a compact semantic embedding, which is then compressed and transmitted for downstream inference. We propose an adaptive transform-coding method for semantic-feature compression motivated by the conditional rate-distortion function of a Gaussian mixture model. The scheme uses mode-dependent transforms and quantizers selected according to the inferred source component, enabling more efficient coding of heterogeneous feature distributions. Evaluations on features from widely used vision backbones and foundation models show that the proposed method outperforms or is competitive with state-of-the-art neural compression methods while preserving flexibility and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。