arXiv:2410.06149cs.CVcs.MM2024-10被引 11

用可适应内容的扩散模型,让图像压缩同时保细节和机器可用性。

Toward Scalable Image Feature Compression: A Content-Adaptive and Diffusion-Based Approach

  • 通过扩散过程编码纹理,融合语义与细节特征
  • 支持无重训练的灵活压缩比调节,重建质量更优
  • 适合需要兼顾人眼感知与机器视觉的压缩场景

传统图像编码侧重信号保真度和人眼感知,常牺牲机器视觉任务性能。深度学习方法虽利用丰富语义嵌入在人机视觉上表现良好,但难以捕捉轮廓、纹理等细粒度信息,导致重建不完美。现有学习型编码器缺乏可扩展性。本文提出一种内容自适应的扩散模型用于可扩展图像压缩:通过马尔可夫调色板扩散模型结合通用特征提取器与图像生成器,实现高效压缩。利用协同纹理-语义特征提取与伪标签生成,准确捕获纹理信息;再以内容自适应的马尔可夫调色板扩散模型,以可扩展方式表示低层纹理与高层语义。通过选择中间扩散状态即可灵活控制压缩率,无需在不同码率下重新训练模型。大量实验表明,该框架在图像重建及目标检测、分割、人脸关键点定位等下游任务中均优于当前最优方法,且感知质量显著提升。

原文摘要 · Abstract (English)

Traditional image codecs emphasize signal fidelity and human perception, often at the expense of machine vision tasks. Deep learning methods have demonstrated promising coding performance by utilizing rich semantic embeddings optimized for both human and machine vision. However, these compact embeddings struggle to capture fine details such as contours and textures, resulting in imperfect reconstructions. Furthermore, existing learning-based codecs lack scalability. To address these limitations, this paper introduces a content-adaptive diffusion model for scalable image compression. The proposed method encodes fine textures through a diffusion process, enhancing perceptual quality while preserving essential features for machine vision tasks. The approach employs a Markov palette diffusion model combined with widely used feature extractors and image generators, enabling efficient data compression. By leveraging collaborative texture-semantic feature extraction and pseudo-label generation, the method accurately captures texture information. A content-adaptive Markov palette diffusion model is then applied to represent both low-level textures and high-level semantic content in a scalable manner. This framework offers flexible control over compression ratios by selecting intermediate diffusion states, eliminating the need for retraining deep learning models at different operating points. Extensive experiments demonstrate the effectiveness of the proposed framework in both image reconstruction and downstream machine vision tasks such as object detection, segmentation, and facial landmark detection, achieving superior perceptual quality compared to state-of-the-art methods.

图像压缩扩散模型内容自适应机器视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。