arXiv:2501.00975cs.CVcs.LG2025-01

CoordFlow通过坐标流实现像素级视频表示,压缩效果媲美顶级帧级方法。

CoordFlow: Coordinate Flow for Pixel-wise Neural Video Representation

  • 将视觉信息分层并用专用网络补偿运动,提升压缩效率。
  • 在像素级方法中达到顶尖性能,与主流帧级方法表现相当。
  • 无需标注即可实现视频分割、超分辨率和去噪等多功能能力。

在视频压缩领域,如何在更低码率下获得更高质量始终是长期目标。近年来,隐式神经表示(INR)展现出替代传统变换编码方法的潜力。视频INR大致可分为基于帧和基于像素的方法。尽管像素级方法在上采样和并行化方面更具优势,但帧级方法表现更优。本文提出CoordFlow,一种新型像素级视频压缩INR,其性能优于其他像素级方法,并与领先帧级技术表现相当。该方法通过将视觉信息分解为视觉一致的多个层次,每层由专用网络建模并补偿其运动。融合后可生成无监督视频分割结果,物体运动轨迹被隐式用于消除时空冗余。此外,该方法天然支持视频超分辨率、稳定化、修复与去噪。

原文摘要 · Abstract (English)

In the field of video compression, the pursuit for better quality at lower bit rates remains a long-lasting goal. Recent developments have demonstrated the potential of Implicit Neural Representation (INR) as a promising alternative to traditional transform-based methodologies. Video INRs can be roughly divided into frame-wise and pixel-wise methods according to the structure the network outputs. While the pixel-based methods are better for upsampling and parallelization, frame-wise methods demonstrated better performance. We introduce CoordFlow, a novel pixel-wise INR for video compression. It yields state-of-the-art results compared to other pixel-wise INRs and on-par performance compared to leading frame-wise techniques. The method is based on the separation of the visual information into visually consistent layers, each represented by a dedicated network that compensates for the layer's motion. When integrated, a byproduct is an unsupervised segmentation of video sequence. Objects motion trajectories are implicitly utilized to compensate for visual-temporal redundancies. Additionally, the proposed method provides inherent video upsampling, stabilization, inpainting, and denoising capabilities.

视频压缩隐式表示像素级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。