arXiv:2511.22549cs.CV2025-11NeurIPS被引 5

用生成先验统一图像压缩中人眼与机器视觉需求

Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative Prior

  • 引入语义一致性损失和标签引导模块,融合生成模型提升压缩质量
  • 单个码流支持多智能任务,且保持人类感知视觉质量
  • 无需任务适配即可兼顾机器理解与人眼体验,适合多场景应用

图像压缩方法通常独立优化人眼感知或机器分析任务。我们发现二者本质共性:保留准确语义信息至关重要,既保障智能任务的关键信息完整性,也促进人类理解。同时,提升感知质量不仅增强视觉吸引力,还能通过维持真实图像分布,改善机器任务的语义特征提取。基于此,我们提出 Diff-ICMH,一种生成式图像压缩框架,通过生成先验保证感知真实感,并在训练中引入语义一致性损失(SC loss)确保语义保真。此外,我们设计标签引导模块(TGM),利用高语义图像级标签激发预训练扩散模型生成能力,仅需极低额外比特率。因此,Diff-ICMH 能通过单一编码器和比特流支持多种智能任务,无需任务特化调整,同时保持高质量人类视觉体验。大量实验表明,该方法在多样化任务中均具优越性与泛化能力。

原文摘要 · Abstract (English)

Image compression methods are usually optimized isolatedly for human perception or machine analysis tasks. We reveal fundamental commonalities between these objectives: preserving accurate semantic information is paramount, as it directly dictates the integrity of critical information for intelligent tasks and aids human understanding. Concurrently, enhanced perceptual quality not only improves visual appeal but also, by ensuring realistic image distributions, benefits semantic feature extraction for machine tasks. Based on this insight, we propose Diff-ICMH, a generative image compression framework aiming for harmonizing machine and human vision in image compression. It ensures perceptual realism by leveraging generative priors and simultaneously guarantees semantic fidelity through the incorporation of Semantic Consistency loss (SC loss) during training. Additionally, we introduce the Tag Guidance Module (TGM) that leverages highly semantic image-level tags to stimulate the pre-trained diffusion model's generative capabilities, requiring minimal additional bit rates. Consequently, Diff-ICMH supports multiple intelligent tasks through a single codec and bitstream without any task-specific adaptation, while preserving high-quality visual experience for human perception. Extensive experimental results demonstrate Diff-ICMH's superiority and generalizability across diverse tasks, while maintaining visual appeal for human perception. Code is available at: https://github.com/RuoyuFeng/Diff-ICMH.

图像压缩生成模型多任务语义保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。