arXiv:2410.10171eess.IVcs.CV2024-10被引 8

用多粒度运动轨迹分解,让人体视频压缩更省带宽且画质更好。

Generative Human Video Compression with Multi-granularity Temporal Trajectory Factorization

  • 通过多粒度运动轨迹分解,将高维视觉信号压缩为紧凑运动向量。
  • 在说话人脸和动态身体视频上,客观与主观质量均优于VVC标准。
  • 支持分辨率自适应,适合低带宽下的人体视频通信场景。

本文提出一种新型的多粒度时间轨迹分解框架,用于生成式人体视频压缩,具有显著降低带宽消耗的潜力。所提出的运动分解策略可将高维视觉信号隐式表征为紧凑的运动向量,实现表示紧凑性;并进一步将其转化为细粒度场,增强运动表达能力。由此,编码码流可在最低表示成本下携带充足的视觉运动信息。同时,设计了具备增强背景稳定性的可扩展分辨率生成模块,使框架在重建鲁棒性和分辨率灵活性方面得到优化。实验结果表明,在说话人脸与动态身体视频上,该方法在客观与主观质量上均优于最新生成模型及最先进的视频编码标准Versatile Video Coding(VVC)。项目主页见:https://github.com/xyzysz/Extreme-Human-Video-Compression-with-MTTF。

原文摘要 · Abstract (English)

In this paper, we propose a novel Multi-granularity Temporal Trajectory Factorization framework for generative human video compression, which holds great potential for bandwidth-constrained human-centric video communication. In particular, the proposed motion factorization strategy can facilitate to implicitly characterize the high-dimensional visual signal into compact motion vectors for representation compactness and further transform these vectors into a fine-grained field for motion expressibility. As such, the coded bit-stream can be entailed with enough visual motion information at the lowest representation cost. Meanwhile, a resolution-expandable generative module is developed with enhanced background stability, such that the proposed framework can be optimized towards higher reconstruction robustness and more flexible resolution adaptation. Experimental results show that proposed method outperforms latest generative models and the state-of-the-art video coding standard Versatile Video Coding (VVC) on both talking-face videos and moving-body videos in terms of both objective and subjective quality. The project page can be found at https://github.com/xyzysz/Extreme-Human-Video-Compression-with-MTTF.

视频压缩生成模型人体视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。