arXiv:2511.08915cs.CV2025-11

用机器视觉主导压缩,实现人机协同的高效图像编码。

Machines Serve Human: A Novel Variable Human-machine Collaborative Compression Framework

  • 以机器视觉为核心构建新压缩框架,替代传统人类视觉流程。
  • 在图像和视频压缩任务中,机器与人类视觉性能均显著提升。
  • 支持可变码率,适合需兼顾精度与效率的应用场景。

人机协同压缩旨在降低图像/视频数据量,服务于人类感知与机器智能。现有方法多基于人类视觉压缩范式,在融合机器视觉压缩时存在复杂度高、码率高的问题。由于机器视觉仅关注图像核心区域,所需信息远少于人类视觉压缩内容。本文首次提出一种面向机器视觉的新型协同压缩方法,将机器视觉作为人机协同压缩的基础。设计了即插即用的可变码率策略以适配机器视觉任务。进一步提出渐进式聚合机器视觉语义特征,并利用扩散先验无缝恢复人类视觉所需的高保真细节,命名为基于扩散先验的跨视觉特征压缩(Diff-FCHM)。实验结果表明,该方法在机器视觉与人类视觉压缩任务上均取得显著更优表现。代码将在论文录用后公开。

原文摘要 · Abstract (English)

Human-machine collaborative compression has been receiving increasing research efforts for reducing image/video data, serving as the basis for both human perception and machine intelligence. Existing collaborative methods are dominantly built upon the de facto human-vision compression pipeline, witnessing deficiency on complexity and bit-rates when aggregating the machine-vision compression. Indeed, machine vision solely focuses on the core regions within the image/video, requiring much less information compared with the compressed information for human vision. In this paper, we thus set out the first successful attempt by a novel collaborative compression method based on the machine-vision-oriented compression, instead of human-vision pipeline. In other words, machine vision serves as the basis for human vision within collaborative compression. A plug-and-play variable bit-rate strategy is also developed for machine vision tasks. Then, we propose to progressively aggregate the semantics from the machine-vision compression, whilst seamlessly tailing the diffusion prior to restore high-fidelity details for human vision, thus named as diffusion-prior based feature compression for human and machine visions (Diff-FCHM). Experimental results verify the consistently superior performances of our Diff-FCHM, on both machine-vision and human-vision compression with remarkable margins. Our code will be released upon acceptance.

人机协同压缩算法扩散模型机器视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。