arXiv:2507.07633cs.CVcs.MM2025-07AAAI被引 9

用轨迹引导生成视频编码,超低码率下仍保细节与连贯性。

T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates

  • 通过语义感知的稀疏运动采样提取关键轨迹点,降低码率
  • 在扩散模型中引入轨迹对齐损失,无需训练即可生成合理运动
  • 适合需要精准运动控制的超低码率视频应用

近期视频生成技术的发展催生了利用强大生成先验进行超低码率(ULB)视频编码的新范式。然而,现有方法受限于领域特定性(如人脸或人体视频)或过度依赖高层文本指导,难以捕捉细粒度运动细节,导致重建结果不真实或不连贯。为此,我们提出轨迹引导生成视频编码(T-GVC),该框架将低层运动跟踪与高层语义理解相结合。T-GVC采用语义感知的稀疏运动采样流程,基于语义重要性提取像素级运动作为稀疏轨迹点,显著降低码率同时保留关键时序语义信息。此外,通过在扩散过程中引入轨迹对齐损失约束,我们实现了无需训练的潜在空间引导机制,确保物理上合理的运动模式,同时不牺牲生成模型的固有能力。实验表明,T-GVC在ULB条件下优于传统及神经视频编码器。进一步实验验证,本框架在运动控制精度上超过现有文本引导方法,为基于几何运动建模的生成视频编码开辟了新方向。

原文摘要 · Abstract (English)

Recent advances in video generation techniques have given rise to an emerging paradigm of generative video coding for Ultra-Low Bitrate (ULB) scenarios by leveraging powerful generative priors. However, most existing methods are limited by domain specificity (e.g., facial or human videos) or excessive dependence on high-level text guidance, which tend to inadequately capture fine-grained motion details, leading to unrealistic or incoherent reconstructions. To address these challenges, we propose Trajectory-Guided Generative Video Coding (dubbed T-GVC), a novel framework that bridges low-level motion tracking with high-level semantic understanding. T-GVC features a semantic-aware sparse motion sampling pipeline that extracts pixel-wise motion as sparse trajectory points based on their semantic importance, significantly reducing the bitrate while preserving critical temporal semantic information. In addition, by integrating trajectory-aligned loss constraints into diffusion processes, we introduce a training-free guidance mechanism in latent space to ensure physically plausible motion patterns without sacrificing the inherent capabilities of generative models. Experimental results demonstrate that T-GVC outperforms both traditional and neural video codecs under ULB conditions. Furthermore, additional experiments confirm that our framework achieves more precise motion control than existing text-guided methods, paving the way for a novel direction of generative video coding guided by geometric motion modeling.

视频编码生成模型轨迹引导超低码率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。