arXiv:2510.10970eess.IV2025-10中稿 · the 2025 Picture C…被引 1

用感知损失训练轻量模型,提升VVC内编码的视觉质量。

Bit Allocation Transfer for Perceptual Quality Enhancement of VVC Intra Coding

  • 通过感知损失训练轻量模型生成量化步长图
  • 在Kodak和CLIC上实现超11%的MS-SSIM BD率降低
  • 适合想快速提升传统编码器视觉质量的研究者

主流图像与视频编码标准(如H.266/VVC、AVS3、AV1)采用基于块的混合编码框架。尽管该框架便于优化峰值信噪比(PSNR),但在优化多尺度结构相似性(MS-SSIM)等感知对齐指标时表现不佳。本文提出一种低复杂度方法,通过将端到端图像压缩中的比特分配知识迁移至VVC内编码,以提升感知质量。引入一个使用感知损失训练的轻量级模型,生成量化步长图,隐式捕捉块级感知重要性,从而高效推导出VVC的QP图。在Kodak和CLIC数据集上的实验表明,该方法在执行时间与感知性能方面均具显著优势,MS-SSIM的BD率降低超过11%。本方案为传统编码器的感知增强提供了高效且实用的路径。

原文摘要 · Abstract (English)

Mainstream image and video coding standards -- including state-of-the-art codecs like H.266/VVC, AVS3, and AV1 -- adopt a block-based hybrid coding framework. While this framework facilitates straightforward optimization for Peak Signal-to-Noise Ratio (PSNR), it struggles to effectively optimize perceptually-aligned metrics such as Multi-Scale Structural Similarity (MS-SSIM). To address this challenge, this paper proposes a low-complexity method to enhance perceptual quality in VVC intra coding by transferring bit allocation knowledge from end-to-end image compression. We introduce a lightweight model trained with perceptual losses to generate a quantization step map. This map implicitly captures block-level perceptual importance, enabling efficient derivation of a QP map for VVC. Experiments on Kodak and CLIC datasets demonstrate significant advantages, both in execution time and perceptual metric performance, with more than 11% BD-rate reduction in terms of MS-SSIM. Our scheme provides an efficient, practical pathway for perceptual enhancement of traditional codecs.

视频编码感知质量比特分配VVC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。