arXiv:2503.21469eess.IVcs.CV2025-03被引 2

让视频压缩失真信息反哺机器视觉,提升分析性能。

Embedding Compression Distortion in Video Coding for Machines

  • 设计特征域压缩敏感提取器,识别机器感知相关的失真
  • 轻量级编码器将失真信息压缩为紧凑表示,传输开销极低
  • 可无缝嵌入下游模型,适合视频分析与智能编码场景

当前视频传输不仅服务于人眼视觉系统(HVS),也支持机器感知分析。然而,现有编码器主要针对像素域和人眼感知指标进行优化,忽视了机器视觉任务的需求。为此,本文提出压缩失真表征嵌入(CDRE)框架,通过提取与机器感知相关的压缩失真表征,并将其嵌入下游模型,以弥补压缩过程中的信息损失,提升任务性能。具体而言,设计了一个压缩敏感特征提取器,在特征空间中识别压缩退化;引入轻量级失真编码器,将失真信息压缩为紧凑表征以降低传输开销;随后逐步将该表征嵌入下游模型,使其能感知压缩退化并改善表现。在多种编码器和下游任务上的实验表明,本框架可在比特率、计算时间与参数量几乎无额外负担的前提下,显著提升现有编码器的率-任务性能。代码与补充材料已公开于 https://github.com/Ws-Syx/CDRE/。

原文摘要 · Abstract (English)

Currently, video transmission serves not only the Human Visual System (HVS) for viewing but also machine perception for analysis. However, existing codecs are primarily optimized for pixel-domain and HVS-perception metrics rather than the needs of machine vision tasks. To address this issue, we propose a Compression Distortion Representation Embedding (CDRE) framework, which extracts machine-perception-related distortion representation and embeds it into downstream models, addressing the information lost during compression and improving task performance. Specifically, to better analyze the machine-perception-related distortion, we design a compression-sensitive extractor that identifies compression degradation in the feature domain. For efficient transmission, a lightweight distortion codec is introduced to compress the distortion information into a compact representation. Subsequently, the representation is progressively embedded into the downstream model, enabling it to be better informed about compression degradation and enhancing performance. Experiments across various codecs and downstream tasks demonstrate that our framework can effectively boost the rate-task performance of existing codecs with minimal overhead in terms of bitrate, execution time, and number of parameters. Our codes and supplementary materials are released in https://github.com/Ws-Syx/CDRE/.

视频编码机器视觉失真建模压缩感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。