提出低复杂度数据自适应视频编码变换,性能接近最优但计算量小。
INT-DTT+: Low-Complexity Data-Dependent Transforms for Video Coding
- 基于图学习和分层结构设计可快速计算的数据依赖变换
- 在VVC标准中实现3%以上码率节省,接近最优变换性能
- 适合追求高效编码的视频压缩系统开发者使用
离散三角变换(DTT),如DCT-2和DST-7,因编码性能与计算效率平衡而广泛用于视频编码。相比之下,数据依赖变换(如KLT和基于图的可分离变换GBSTs)虽具更好能量集中性,却缺乏可利用的对称性以降低计算复杂度。本文提出通用框架,设计低复杂度数据依赖变换。方法基于DTT+,即通过秩一更新构建的DTT图族,能适应信号统计特性并保留快速计算结构。我们提出联合估计行列图秩一更新的图学习算法,捕捉块的整体统计特征;利用DTT+的渐进结构,将核分解为基DTT与结构化Cauchy矩阵。结合低复杂度整数DTT与稀疏化Cauchy矩阵,构造出整数近似变换INT-DTT+,显著降低计算与内存开销,性能损失极小。在符合率失真优化(RDOT)设计的模式依赖变换场景下验证,集成至VVC的显式多变换选择(MTS)框架后,相比VVC MTS基准,INT-DTT+实现超过3%的BD-rate节省,且复杂度与整数DCT-2相当(基DTT系数可用时)。
原文摘要 · Abstract (English)
Discrete trigonometric transforms (DTTs), such as the DCT-2 and the DST-7, are widely used in video codecs for their balance between coding performance and computational efficiency. In contrast, data-dependent transforms, such as the Karhunen-Loève transform (KLT) and graph-based separable transforms (GBSTs), offer better energy compaction but lack symmetries that can be exploited to reduce computational complexity. This paper bridges this gap by introducing a general framework to design low-complexity data-dependent transforms. Our approach builds on DTT+, a family of GBSTs derived from rank-one updates of the DTT graphs, which can adapt to signal statistics while retaining a structure amenable to fast computation. We first propose a graph learning algorithm for DTT+ that estimates the rank-one updates for rows and column graphs jointly, capturing the statistical properties of the overall block. Then, we exploit the progressive structure of DTT+ to decompose the kernel into a base DTT and a structured Cauchy matrix. By leveraging low-complexity integer DTTs and sparsifying the Cauchy matrix, we construct an integer approximation to DTT+, termed INT-DTT+. This approximation significantly reduces both computational and memory complexities with respect to the separable KLT with minimal performance loss. We validate our approach in the context of mode-dependent transforms for the VVC standard, following a rate-distortion optimized transform (RDOT) design approach. Integrated into the explicit multiple transform selection (MTS) framework of VVC in a rate-distortion optimization setup, INT-DTT+ achieves more than 3% BD-rate savings over the VVC MTS baseline, with complexity comparable to the integer DCT-2 once the base DTT coefficients are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。