CoordFlow通过坐标流实现像素级视频表示,压缩效果媲美顶级帧级方法。
CoordFlow: Coordinate Flow for Pixel-wise Neural Video Representation
- 将视觉信息分层并用专用网络补偿运动,提升压缩效率。
- 在像素级方法中达到顶尖性能,与主流帧级方法表现相当。
- 无需标注即可实现视频分割、超分辨率和去噪等多功能能力。
在视频压缩领域,如何在更低码率下获得更高质量始终是长期目标。近年来,隐式神经表示(INR)展现出替代传统变换编码方法的潜力。视频INR大致可分为基于帧和基于像素的方法。尽管像素级方法在上采样和并行化方面更具优势,但帧级方法表现更优。本文提出CoordFlow,一种新型像素级视频压缩INR,其性能优于其他像素级方法,并与领先帧级技术表现相当。该方法通过将视觉信息分解为视觉一致的多个层次,每层由专用网络建模并补偿其运动。融合后可生成无监督视频分割结果,物体运动轨迹被隐式用于消除时空冗余。此外,该方法天然支持视频超分辨率、稳定化、修复与去噪。
原文摘要 · Abstract (English)
In the field of video compression, the pursuit for better quality at lower bit rates remains a long-lasting goal. Recent developments have demonstrated the potential of Implicit Neural Representation (INR) as a promising alternative to traditional transform-based methodologies. Video INRs can be roughly divided into frame-wise and pixel-wise methods according to the structure the network outputs. While the pixel-based methods are better for upsampling and parallelization, frame-wise methods demonstrated better performance. We introduce CoordFlow, a novel pixel-wise INR for video compression. It yields state-of-the-art results compared to other pixel-wise INRs and on-par performance compared to leading frame-wise techniques. The method is based on the separation of the visual information into visually consistent layers, each represented by a dedicated network that compensates for the layer's motion. When integrated, a byproduct is an unsupervised segmentation of video sequence. Objects motion trajectories are implicitly utilized to compensate for visual-temporal redundancies. Additionally, the proposed method provides inherent video upsampling, stabilization, inpainting, and denoising capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。