将视频运动分层编码,提升压缩效率与细节保留
MotionStrata: Hierarchical Motion Latents for Compact Video Autoencoding
- 分全局运动与细节运动两层编码,共享固定码本预算
- 在极端压缩下仍保持高重建质量,优于均匀或分组编码
- 适合需要高效视频压缩与生成的场景
首帧条件视频自编码器通过持久内容和紧凑运动码本表示视频片段。尽管已消除大量外观冗余,剩余运动通常采用同质潜空间压缩,对整体场景演化和帧级细节使用相同时间支持。本文提出MotionStrata,将固定运动预算划分为全局运动与细节运动两层:时序压缩的全局查询捕捉整体演化,帧对齐的细节查询保留随帧变化的精细结构。频率引导路由与自粗到精训练构建该层次结构,不增加运动维度。实验表明,MotionStrata在激进压缩下维持高重建质量,优于均匀及其它分组表示。额外实验评估了层次化表示、下游生成与解码开销,结果支持分层运动组织作为紧凑视频自编码的有效设计原则。
原文摘要 · Abstract (English)
First-frame-conditioned video autoencoders represent a clip with persistent content and a compact motion code. Although this removes much of the appearance redundancy, the remaining motion is typically compressed with a homogeneous latent geometry. Such representations use the same temporal support for broad scene evolution and fine-grained, frame-specific details. We introduce MotionStrata, which organizes a fixed motion budget into Global Motion and Detailed Motion. Temporally compressed Global queries summarize broad evolution, whereas frame-aligned Detailed queries preserve fine-grained structures whose configuration varies across frames. Frequency-guided routing and coarse-to-fine training establish this hierarchy without increasing motion dimensionality. Experiments show that MotionStrata maintains high reconstruction quality under aggressive compression and outperforms uniform and alternative grouped representations. Additional experiments evaluate hierarchical representation, downstream generation, and decoding cost. These results support hierarchical motion organization as a useful design principle for compact video autoencoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。