arXiv:2512.10450cs.CV2025-12

解决视频压缩中运动估计误差传播问题,实现无误差累积的高质量流式传输。

Error-Propagation-Free Learned Video Compression With Dual-Domain Progressive Temporal Alignment

  • 采用双域渐进对齐机制,先像素域粗对齐再隐空间精对齐,提升运动建模能力。
  • 在多个数据集上达到领先水平的码率-失真性能,且完全消除误差传播。
  • 支持连续码率自适应,适合需要稳定画质的实时视频流场景。

现有学习型视频压缩框架在运动估计与补偿(ME/MC)中面临对齐不准与误差传播的两难困境。分离变换框架虽有优异码率-失真(R-D)表现,但存在明显误差传播;统一变换框架通过共享变换消除误差传播,却在共享隐空间中运动建模能力较弱。为此,本文提出一种新型统一变换框架,结合双域渐进时间对齐与质量条件混合专家(QCMoE),实现质量一致且无误差传播的流式压缩。具体而言,双域渐进对齐通过粗粒度像素域对齐与精细隐空间对齐,以粗到精方式增强时序上下文建模:像素域对齐利用单参考帧光流处理简单运动,隐空间对齐则基于多参考帧隐表示引入流引导可变形变压器(FGDT),实现复杂运动的长期运动精修(LTMR)。此外,设计了QCMoE模块用于连续码率自适应,动态按目标质量与内容分配不同专家调整每像素量化步长,而非依赖单一量化步长。该模块实现平滑一致的码率控制并取得优异的R-D性能。实验表明,所提方法在多个基准上达到先进水平,同时成功消除误差传播。

原文摘要 · Abstract (English)

Existing frameworks for learned video compression suffer from a dilemma between inaccurate temporal alignment and error propagation for motion estimation and compensation (ME/MC). The separate-transform framework employs distinct transforms for intra-frame and inter-frame compression to yield impressive rate-distortion (R-D) performance but causes evident error propagation, while the unified-transform framework eliminates error propagation via shared transforms but is inferior in ME/MC in shared latent domains. To address this limitation, in this paper, we propose a novel unifiedtransform framework with dual-domain progressive temporal alignment and quality-conditioned mixture-of-expert (QCMoE) to enable quality-consistent and error-propagation-free streaming for learned video compression. Specifically, we propose dualdomain progressive temporal alignment for ME/MC that leverages coarse pixel-domain alignment and refined latent-domain alignment to significantly enhance temporal context modeling in a coarse-to-fine fashion. The coarse pixel-domain alignment efficiently handles simple motion patterns with optical flow estimated from a single reference frame, while the refined latent-domain alignment develops a Flow-Guided Deformable Transformer (FGDT) over latents from multiple reference frames to achieve long-term motion refinement (LTMR) for complex motion patterns. Furthermore, we design a QCMoE module for continuous bit-rate adaptation that dynamically assigns different experts to adjust quantization steps per pixel based on target quality and content rather than relies on a single quantization step. QCMoE allows continuous and consistent rate control with appealing R-D performance. Experimental results show that the proposed method achieves competitive R-D performance compared with the state-of-the-arts, while successfully eliminating error propagation.

视频压缩运动估计无误差传播自适应量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。