arXiv:2602.15779eess.IV2026-02

用多非参考度量优化视频编码,提升质量稳定性与效率

Rate-Distortion Optimization for Ensembles of Non-Reference Metrics

  • 通过集成多个非参考度量并平滑梯度,改进编码器决策
  • 在多个度量上实现稳定码率节省,且编码速度大幅提升
  • 适合过拟合编码器,无需反向传播神经网络度量

非参考度量(NRMs)可在无参考图像的情况下评估视频质量,特别适用于用户生成内容的评价。然而,当前视频编码中的率失真优化(RDO)仍主要依赖全参考度量(如均方误差),将输入视为理想目标。将NRMs引入RDO的一种方法是线性化(LNRM),即利用NRM对输入的梯度指导比特分配。但该方法在某些其他NRMs上表现有限甚至退化。本文指出,NRMs是高度非线性的预测器,其梯度局部不稳定,会损害线性化效果;且优化单一度量可能引入模型特有偏差,难以泛化。为此,本文扩展了LNRM框架,以优化多个NRMs的集合,并引入基于平滑的公式,在线性化前稳定梯度。该方法适用于混合编码器,尤其适合过拟合编码器,避免对神经网络型NRMs进行迭代评估和反向传播,显著降低编码复杂度。在AVC与Cool-chic编码器上使用YouTube UGC数据集验证,实验显示在多个NRMs上均实现一致的码率节省,无解码器开销,且在Cool-chic上编码时间大幅减少。

原文摘要 · Abstract (English)

Non-reference metrics (NRMs) can assess the visual quality of images and videos without a reference, making them well-suited for the evaluation of user-generated content. Nonetheless, rate-distortion optimization (RDO) in video coding is still mainly driven by full-reference metrics, such as the sum of squared errors, which treat the input as an ideal target. A way to incorporate NRMs into RDO is through linearization (LNRM), where the gradient of the NRM with respect to the input guides bit allocation. While this strategy improves the quality predicted by some metrics, we show that it can yield limited gains or degradations when evaluated with other NRMs. We argue that NRMs are highly non-linear predictors with locally unstable gradients that can compromise the quality of the linearization; furthermore, optimizing a single metric may exploit model-specific biases that do not generalize across quality estimators. Motivated by this observation, we extend the LNRM framework to optimize ensembles of NRMs and, to further improve robustness, we introduce a smoothing-based formulation that stabilizes NRM gradients prior to linearization. Our framework is well-suited to hybrid codecs, and we advocate for its use with overfitted codecs, where it avoids iterative evaluations and backpropagation of neural network-based NRMs, reducing encoder complexity relative to direct NRM optimization. We validate the proposed approach on AVC and Cool-chic, using the YouTube UGC dataset. Experiments demonstrate consistent bitrate savings across multiple NRMs with no decoder complexity overhead and, for Cool-chic, a substantial reduction in encoding runtime compared to direct NRM optimization.

视频编码非参考度量率失真优化编码效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。