用可微代理实现自适应量化,提升H.264视频编码的感知质量与任务性能。
Differentiable Proxy Learning for Adaptive Quantization Control in H.264 Video Coding

- 构建可微代理模型,通过软索引机制实现对编码器QP的梯度优化。
- 在全局和空间级量化下逼近H.264的率失真特性,实测BD-rate降低最高17.12%。
- 适用于感知优化与机器视觉任务,显著优于固定QP的基准方案。
H.264因其简洁性、高效性及广泛软硬件支持,过去二十年间成为最主流的视频编码格式。然而,由于标准编解码器不可微,针对特定目标(如感知质量或机器视觉任务)优化量化参数(QP)极具挑战。尽管已有可微代理用于实现编解码器的梯度优化,但其对目标编解码器的拟合精度常未明确评估。本文提出一种针对H.264帧内编码的可微代理学习方法,以实现自适应量化控制。该代理基于可变率学习压缩模型,通过软索引机制使其对编解码器QP可微。代理在两种量化设置下训练:全局QP(每帧一个QP)与空间级QP(宏块级分配)。使用冻结训练后的代理,我们构建了面向感知优化与机器视觉任务的代理式自适应量化(AQ)框架。实验表明,所提代理能紧密逼近H.264帧内编码的率失真行为。由此产生的代理式AQ框架在多个任务上持续改进率-任务权衡,相较于固定QP H.264基线,在语义分割任务中实现最高达17.12%的BD-rate降低,于MS-SSIM任务中实现15.30%的降低。
原文摘要 · Abstract (English)
H.264 has been the most widely used video coding format for the past two decades due to its relative simplicity, efficiency, and wide availability of software and hardware implementations. However, optimizing codec parameters such as the quantization parameter (QP) for specific objectives (e.g., perceptual quality or machine vision tasks) is challenging due to the non-differentiable nature of standard video codecs. While differentiable proxies have recently been used to enable gradient-based optimization around standard codecs, their fidelity to the target codec is rarely explicitly characterized. In this paper, we propose a differentiable proxy learning method for H.264 intra codec to enable adaptive quantization control. Built upon a variable-rate learned compression model, the proposed proxy is made differentiable with respect to codec QP through a soft-indexing mechanism. It is then trained to approximate the rate-distortion behavior of H.264 under two quantization settings: global-QP, which uses one QP per image, and spatial-QP, which assigns QPs at the macroblock level. Using the frozen trained proxy, we develop a proxy-based adaptive quantization (AQ) framework for both perceptual optimization and machine vision tasks. Experimental results demonstrate that the proposed proxies closely approximate the rate-distortion behavior of H.264 intra codec. The resulting proxy-based AQ framework consistently improves rate-task trade-offs over fixed-QP H.264 baselines, achieving BD-rate reduction of up to 17.12% for semantic segmentation and 15.30% for MS-SSIM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。