arXiv:2512.12936cs.CVcs.AI2025-12中稿 · Data Compression C…被引 6

提出自适应运动对齐框架,提升视频压缩效率与质量。

Content Adaptive based Motion Alignment Framework for Learned Video Compression

  • 分两阶段预测运动偏移并调节掩码,实现精细特征对齐。
  • 根据参考帧质量动态调整失真权重,减少误差传播。
  • 无需训练的下采样模块,按运动幅度与分辨率优化估计。

端到端视频压缩虽取得进展,但普遍缺乏内容自适应能力,导致压缩性能受限。本文提出基于内容自适应的运动对齐框架(CAMA),通过三方面改进提升性能:首先设计两阶段流引导可变形扭曲机制,通过粗到细的偏移预测与掩码调制,实现精准特征对齐;其次提出多参考质量感知策略,依据参考帧质量动态调整失真权重,并应用于分层训练以抑制误差传播;最后引入无需训练的下采样模块,根据运动幅度和分辨率对帧进行降采样,获得更平滑的运动估计。在标准测试数据集上的实验表明,CAMA 在对比基线模型 DCVC-TCM 时,实现了 24.95% 的 BD-rate(PSNR)节省,同时优于复现的 DCVC-DC 和传统编码器 HM-16.25。

原文摘要 · Abstract (English)

Recent advances in end-to-end video compression have shown promising results owing to their unified end-to-end learning optimization. However, such generalized frameworks often lack content-specific adaptation, leading to suboptimal compression performance. To address this, this paper proposes a content adaptive based motion alignment framework that improves performance by adapting encoding strategies to diverse content characteristics. Specifically, we first introduce a two-stage flow-guided deformable warping mechanism that refines motion compensation with coarse-to-fine offset prediction and mask modulation, enabling precise feature alignment. Second, we propose a multi-reference quality aware strategy that adjusts distortion weights based on reference quality, and applies it to hierarchical training to reduce error propagation. Third, we integrate a training-free module that downsamples frames by motion magnitude and resolution to obtain smooth motion estimation. Experimental results on standard test datasets demonstrate that our framework CAMA achieves significant improvements over state-of-the-art Neural Video Compression models, achieving a 24.95% BD-rate (PSNR) savings over our baseline model DCVC-TCM, while also outperforming reproduced DCVC-DC and traditional codec HM-16.25.

视频压缩运动对齐自适应编码神经网络

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。