提出细粒度退化引导的视频增强方法,无需已知压缩参数仍可显著提升画质。
Blind Quality Enhancement of Compressed Video via Fine-Grained Degradation-Guided Sequential Inference
- 从压缩视频中解耦提取多尺度退化特征,提供空间精细的修复指引。
- 在QP=22时,PSNR提升达0.65dB,较之前最佳盲法提升0.34dB。
- 自适应分阶段推理,使平均推理时间减少50%,适合实际部署。
现有压缩视频画质增强(QECV)方法大多依赖已知量化参数(QPs),需为每种参数训练独立模型,属于非盲方法。但在转码和传输等实际场景中,QPs可能部分或完全不可用,限制了其应用,推动了盲式QECV技术的发展。现有盲法通常使用交叉熵训练的分类模型生成退化向量,作为通道注意力引导去噪,但此类表示主要捕捉全局压缩信息,缺乏细粒度空间线索,难以应对空间变化的伪影模式。为此,我们提出一个预训练退化表征学习模块,能解耦并提取高维、多尺度的退化表征,为伪影消除提供精细引导。此外,多数现有盲法与非盲法采用统一推理架构,忽略不同QP下的计算需求差异。为此,我们引入一种序列化推理策略,根据估计的压缩等级自适应调整伪影消除阶段数。大量实验表明,所提方法显著提升增强性能:在QP=22时,相比当前最优盲法,PSNR提升由0.31dB增至0.65dB;同时,该策略使QP=22下的平均推理时间相比QP=42降低50%。
原文摘要 · Abstract (English)
Existing studies on quality enhancement for compressed video (QECV) predominantly rely on known quantization parameters (QPs), training separate enhancement models for each QP setting, which are referred to as non-blind methods. However, in practical scenarios such as transcoding and transmission, QPs may be partially or entirely unavailable, which limits the applicability of these methods and motivates the development of blind QECV techniques. Existing blind methods typically generate degradation vectors using classification models trained with cross-entropy loss, and employ them as channel attention to guide artifact reduction. Nevertheless, such degradation representations mainly capture global compression information and lack fine-grained spatial cues, making them less effective in handling spatially varying artifact patterns. To address this issue, we propose a pre-trained degradation representation learning module that decouples and extracts high-dimensional, multi-scale degradation representations from compressed video content, providing fine-grained guidance for artifact reduction. Furthermore, most existing blind and nonblind methods adopt a uniform inference architecture for all compression levels, ignoring the distinct computational demands of different QPs. To overcome this limitation, we introduce a sequential inference strategy that adaptively adjusts the number of artifact reduction stages according to the estimated compression level. Extensive experiments show that the proposed method significantly improves enhancement performance. In particular, at QP = 22, it raises PSNR improvement from 0.31 dB to 0.65 dB over the previous state-of-the-art blind method. Meanwhile, with the proposed sequential inference strategy, the average inference time at QP = 22 is reduced by 50% compared with that at QP = 42.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。