arXiv:2606.13957eess.IVcs.CV2026-06

首个单架构覆盖全码率范围的高保真神经视频压缩方法

High-Fidelity Video Compression based on Invertible Neural Transform and Implicit Conditioning

论文配图:High-Fidelity Video Compression based on Invertible Neural Transform and Implicit Conditioning
图 1 · 摘自论文原文
  • 采用可逆主变换+隐式条件场,分离内容与细节以提升重建效率
  • 在UVG数据集上实现PSNR提升21.66%、MS-SSIM提升46.06%的码率节省
  • 适用于从低码率到高保真全范围压缩,尤其适合高质量场景

基于学习的视频压缩近年在速率-失真性能上已接近传统编码器。但多数方法依赖非可逆的分析-合成变换,重建质量受量化误差和变换近似误差双重影响。这一限制在高保真场景尤为突出,此时量化误差小,变换误差成为主导。为此,我们提出InnVC——一种基于可逆神经网络的视频编码器,支持宽范围与高保真压缩。核心思想是在量化前保持可逆主变换路径,并通过紧凑的隐式条件场注入内容自适应上下文。该设计将强相关视频内容与难建模的细微细节解耦,使各组件专精于互补重建任务,提升压缩效率。为进一步优化可压缩性,引入调度掩码策略,逐步将信息集中到更少的潜在通道中,增强熵编码效果。在UVG和MCL-JCV基准测试中,InnVC在广泛质量范围内表现优异,尤其在高保真段,相比x265在UVG上实现21.66%的BD-rate降低(PSNR)和46.06%(MS-SSIM)。据我们所知,InnVC是首个在单一架构内覆盖从低码率至高保真全范围运行点的神经视频编码器,涵盖超过20 dB的PSNR动态范围。

原文摘要 · Abstract (English)

Learning-based video compression has recently achieved competitive rate-distortion performance compared to conventional video codecs. However, most existing methods rely on non-invertible analysis-synthesis transforms, with reconstruction quality subject to both quantization and transform approximation errors. This limitation becomes particularly restrictive at higher quality points, where quantization errors are small and transform-induced distortion dominates. To address this, we propose InnVC, an Invertible neural network based Video Codec for wide-range and high-fidelity compression. The core idea is to preserve an invertible main transform path prior to quantization, while injecting content-adaptive context through a compact implicit conditioning field. This decouples strongly correlated video content from harder-to-model fine details, allowing different components to specialize in complementary reconstruction tasks for more efficient compression. To further improve compressibility, we introduce a scheduled masking strategy that progressively concentrates informative content into fewer latent channels for more effective entropy coding. Experiments on the UVG and MCL-JCV benchmarks show that InnVC achieves strong compression performance over a broad quality range, being particularly effective in the high-quality regime, yielding BD-rate reductions of 21.66% in PSNR and 46.06% in MS-SSIM relative to x265 on UVG. To the best of our knowledge, InnVC is the first neural video codec covers operating poins from low bitrate to high fidelity within a single architecture scale, spanning more than 20 dB in PSNR.

视频压缩可逆网络高保真神经编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。