arXiv:2503.17907cs.CVeess.IV2025-03被引 2

用引导扩散模型实现机器与人眼视觉的无缝图像压缩转换。

Guided Diffusion for the Extension of Machine Vision to Human Visual Perception

  • 用ICM输出引导扩散模型生成人眼可感知图像。
  • 在零额外码率开销下实现机器与人眼视觉的图像重建。
  • 适合需要兼顾机器识别与人类观看的跨模态图像编码场景。

图像压缩技术通过消除冗余信息,实现图像的高效传输与存储,同时服务于机器视觉与人类视觉感知。长期以来,面向人类感知的图像编码已得到深入研究,并催生了多种图像压缩标准。然而,随着图像识别模型的快速发展,面向人工智能任务的图像压缩(即图像编码用于机器,ICM)日益重要。因此,能够同时满足机器与人类需求的可扩展图像编码技术成为研究热点。此外,扩散模型因其能从少量数据生成人眼可感知图像,正被越来越多地应用于人类视觉的图像压缩中。利用扩散模型并以少量条件信息引导生成过程,可部分重构目标图像。受此启发,本文提出一种基于引导扩散的方法,将机器视觉扩展至人类视觉感知。通过使用ICM输出作为引导,从随机噪声中生成人眼可感知图像,引导扩散模型充当机器与人类视觉之间的桥梁,实现二者间无额外比特率开销的转换。实验评估了生成图像的码率与图像质量,结果表明该方法在人类与机器的可扩展图像编码性能上优于现有方案。

原文摘要 · Abstract (English)

Image compression technology eliminates redundant information to enable efficient transmission and storage of images, serving both machine vision and human visual perception. For years, image coding focused on human perception has been well-studied, leading to the development of various image compression standards. On the other hand, with the rapid advancements in image recognition models, image compression for AI tasks, known as Image Coding for Machines (ICM), has gained significant importance. Therefore, scalable image coding techniques that address the needs of both machines and humans have become a key area of interest. Additionally, there is increasing demand for research applying the diffusion model, which can generate human-viewable images from a small amount of data to image compression methods for human vision. Image compression methods that use diffusion models can partially reconstruct the target image by guiding the generation process with a small amount of conditioning information. Inspired by the diffusion model's potential, we propose a method for extending machine vision to human visual perception using guided diffusion. Utilizing the diffusion model guided by the output of the ICM method, we generate images for human perception from random noise. Guided diffusion acts as a bridge between machine vision and human vision, enabling transitions between them without any additional bitrate overhead. The generated images then evaluated based on bitrate and image quality, and we compare their compression performance with other scalable image coding methods for humans and machines.

扩散模型图像压缩机器视觉人眼感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。