arXiv:2509.05929eess.IV2025-09被引 1

用三维率失真复杂度分析,帮用户选最适合的神经视频编码器。

Application Space and the Rate-Distortion-Complexity Analysis of Neural Video CODECs

  • 提出在三维空间中评估编码器的率失真复杂度,用拉格朗日成本统一衡量。
  • 发现只有四个神经视频编码器在不同应用场景下表现最优。
  • 适合需要权衡压缩率、画质和计算成本的开发者参考。

本文通过率-失真-复杂度(RDC)分析,研究视频压缩系统的选择决策过程。讨论了二维Bjontegaard delta(BD)度量,并尝试将其推广至三维RDC空间。进一步探讨了在RDC空间中度量的计算方法,以及如何定义和衡量编码解码器对(codec pair)的成本,其中编码器表现为RDC空间中的点云。采用拉格朗日成本函数 $D+λR + γC$,选择最优编码器需确定合适的 $(λ, γ)$ 值。因此,我们提出每个应用场景可对应 $(λ, γ)$ 平面中的一个点。以流媒体应用为例,在 $(λ, γ)$ 平面中定位特定点。结果表明,可在RDC空间中比较不同编码器的拉格朗日成本。同时,可通过遍历平面,比较所有 $(λ, γ)$ 组合下的编码器性能。我们对比了多个前沿神经视频编码器,结果具启发性且出人意料:在给定的RDC计算约束下,仅有四个神经视频编码器在不同应用需求下表现最佳,具体取决于其期望的 $(λ, γ)$ 位置。

原文摘要 · Abstract (English)

We study the decision-making process for choosing video compression systems through a rate-distortion-complexity (RDC) analysis. We discuss the 2D Bjontegaard delta (BD) metric and formulate generalizations in an attempt to extend its notions to the 3D RDC volume. We follow that discussion with another one on the computation of metrics in the RDC volume, and on how to define and measure the cost of a coder-decoder (codec) pair, where the codec is characterized by a cloud of points in the RDC space. We use a Lagrangian cost $D+λR + γC$, such that choosing the best video codec among a number of candidates for an application demands selecting appropriate $(λ, γ)$ values. Thus, we argue that an application may be associated with a $(λ, γ)$ point in the application space. An example streaming application was given as a case study to set a particular point in the $(λ, γ)$ plane. The result is that we can compare Lagrangian costs in an RDC volume for different codecs for a given application. Furthermore, we can span the plane and compare codecs for the entire application space filled with different $(λ, γ)$ choices. We then compared several state-of-the-art neural video codecs using the proposed metrics. Results are informative and surprising. We found that, within our RDC computation constraints, only four neural video codecs came out as the best suited for any application, depending on where its desirable $(λ, γ)$ lies.

视频编码率失真AI压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。