arXiv:2510.07283eess.IV2025-10中稿 · publication in the…被引 6

通过自适应下采样控制运动矢量范围,提升复杂视频的压缩效率。

Content-Adaptive Inference for State-of-the-art Learned Video Compression

  • 推理时动态下采样帧以匹配训练数据的运动范围。
  • 在单个视频上将DCVC-FM的BD-rate提升最高达41%。
  • 无需微调模型,适用于复杂运动场景的视频压缩。

尽管近期基于学习的视频编码模型在低延迟和随机访问模式下的BD-rate性能平均优于传统编码器,但在具有复杂/大运动的视频上性能提升较小,原因在于学习编码器对训练集中未见的运动矢量范围泛化能力差,导致光流场编码与帧预测性能下降。为此,我们提出一种通用(模型无关)的推理框架,在编码过程中自适应地调整场景中运动矢量的尺度,通过动态下采样帧使测试视频的运动范围近似匹配训练数据。该方法使下采样的运动矢量实现:i)更优的光流估计,从而改善帧预测;ii)更高效的光流压缩。实验表明,该内容自适应推理框架可使现有顶尖低延迟编码器DCVC-FM在单个视频上的BD-rate性能提升最高达41%,且无需模型微调。消融实验验证了运动强度与场景复杂度指标可用于预测该框架的有效性。

原文摘要 · Abstract (English)

While the BD-rate performance of recent learned video codec models in both low-delay and random-access modes exceed that of respective modes of traditional codecs on average over common benchmarks, the performance improvements for individual videos with complex/large motions is much smaller compared to scenes with simple motion. This is related to the inability of a learned encoder model to generalize to motion vector ranges that have not been seen in the training set, which causes loss of performance in both coding of flow fields as well as frame prediction and coding. As a remedy, we propose a generic (model-agnostic) framework to control the scale of motion vectors in a scene during inference (encoding) to approximately match the range of motion vectors in the test and training videos by adaptively downsampling frames. This results in down-scaled motion vectors enabling: i) better flow estimation; hence, frame prediction and ii) more efficient flow compression. We show that the proposed framework for content-adaptive inference improves the BD-rate performance of already state-of-the-art low-delay video codec DCVC-FM by up to 41\% on individual videos without any model fine tuning. We present ablation studies to show measures of motion and scene complexity can be used to predict the effectiveness of the proposed framework.

视频压缩自适应推理光流估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。