arXiv:2609.03520cs.CVcs.AI2026-09

用可变形对齐与差异感知融合提升视频压缩时序信息准确性

Neural Video Compression Based on Deformable Temporal Alignment and Difference-aware Fusion

  • 通过可变形时序对齐生成互补时序上下文
  • 差异感知融合模块自适应选择可靠信息并抑制错位
  • 在复杂运动场景下显著改善压缩质量,适合视频编解码研究者

在基于条件编码的神经视频压缩中,时序上下文质量直接影响压缩性能。现有方法多依赖传播的参考特征构建上下文,但在复杂运动、遮挡和高频纹理区域易受运动估计与局部对齐误差影响,导致时序信息不准确。为此,本文提出结合可变形时序对齐与差异感知空间选择性融合的方法。通过上下文感知时序对齐模块生成互补时序上下文,利用差异感知空间选择性融合模块自适应选择可靠时序信息并抑制错位。实验表明,该方法在率失真性能上优于DCVC-DC。

原文摘要 · Abstract (English)

In conditional coding-based neural video compression, the quality of temporal context directly affects compression per- formance. Existing methods mostly construct context from prop- agated reference features, but they are vulnerable to motion esti- mation and local alignment errors in regions with complex mo- tion, occlusion, and high-frequency textures, resulting in inaccu- rate temporal information. To address this issue, this paper pro- poses a method combining deformable temporal alignment and difference-aware spatial selective fusion. A Context-aware Tem- poral Alignment Module is used to generate complementary tem- poral context, while a Difference-aware Spatial Selective Fusion module adaptively selects reliable temporal information and sup- presses misalignment. Experiments show that the proposed method achieves certain rate-distortion performance improve- ment over DCVC-DC.

视频压缩可变形对齐神经编码时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。