arXiv:2412.08912cs.CVcs.MM2024-12被引 6

首个针对压缩伪影的8K视频修复模型,不依赖噪声假设。

Reversing the Damage: A QP-Aware Transformer-Diffusion Approach for 8K Video Restoration under Codec Compression

  • 用Transformer-扩散模型直接逆转编码器造成的复杂伪影
  • 在4K/8K视频上超越现有方法,尤其在高分辨率下表现突出
  • 适合需要高质量视频还原的影视、广电领域应用

本文提出DiQP,一种用于恢复被编码器压缩(如AV1、HEVC)导致质量下降的8K视频的新型Transformer-扩散模型。据我们所知,该模型是首个在不引入额外噪声的前提下,通过去噪扩散来建模多种编码器带来的复杂非高斯伪影的方法。该设计使模型能够有效学习逆向退化过程。架构结合了Transformer对长程依赖的捕捉能力与增强的窗口机制,以保持跨帧像素组的时空上下文。为提升修复效果,模型还引入了“前瞻”和“环顾”辅助模块,分别提供未来帧与邻近帧信息,有助于重建细节并提升整体视觉质量。在多个数据集上的大量实验表明,该模型在4K和8K等高分辨率视频上均优于当前最优方法,展现出从高度压缩源中恢复感知质量优异视频的有效性。

原文摘要 · Abstract (English)

In this paper, we introduce DiQP; a novel Transformer-Diffusion model for restoring 8K video quality degraded by codec compression. To the best of our knowledge, our model is the first to consider restoring the artifacts introduced by various codecs (AV1, HEVC) by Denoising Diffusion without considering additional noise. This approach allows us to model the complex, non-Gaussian nature of compression artifacts, effectively learning to reverse the degradation. Our architecture combines the power of Transformers to capture long-range dependencies with an enhanced windowed mechanism that preserves spatiotemporal context within groups of pixels across frames. To further enhance restoration, the model incorporates auxiliary "Look Ahead" and "Look Around" modules, providing both future and surrounding frame information to aid in reconstructing fine details and enhancing overall visual quality. Extensive experiments on different datasets demonstrate that our model outperforms state-of-the-art methods, particularly for high-resolution videos such as 4K and 8K, showcasing its effectiveness in restoring perceptually pleasing videos from highly compressed sources.

视频修复扩散模型8K视频编码伪影

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。