arXiv:2509.08376cs.CV2025-09ICCV被引 2

用可控码率的扩散模型分离视频运动与内容,实现自监督学习。

Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video

  • 基于Transformer架构,联合生成帧级运动与片段级内容特征。
  • 通过低码率向量量化构建信息瓶颈,形成有意义的离散运动空间。
  • 适用于真人说话头像、2D动画角色等多类视频,适合视频生成与动作迁移任务。

我们提出一种新颖且通用的框架,将视频数据解耦为动态运动和静态内容两部分。该方法为自监督流水线,相比以往工作假设更少、归纳偏置更低:利用基于Transformer的架构,联合生成帧级运动与片段级内容的灵活隐式特征,并引入低码率向量量化作为信息瓶颈,促进解耦并构建有意义的离散运动空间。以可控码率的潜在运动与内容作为条件输入,驱动去噪扩散模型,实现自监督表征学习。我们在真实世界说话头像视频上验证了该方法在动作迁移与自回归运动生成任务中的有效性;此外,还展示了其在2D卡通角色像素精灵等其他视频类型上的泛化能力。本工作为自监督解耦视频表征学习提供了新视角,推动视频分析与生成领域的发展。

原文摘要 · Abstract (English)

We propose a novel and general framework to disentangle video data into its dynamic motion and static content components. Our proposed method is a self-supervised pipeline with less assumptions and inductive biases than previous works: it utilizes a transformer-based architecture to jointly generate flexible implicit features for frame-wise motion and clip-wise content, and incorporates a low-bitrate vector quantization as an information bottleneck to promote disentanglement and form a meaningful discrete motion space. The bitrate-controlled latent motion and content are used as conditional inputs to a denoising diffusion model to facilitate self-supervised representation learning. We validate our disentangled representation learning framework on real-world talking head videos with motion transfer and auto-regressive motion generation tasks. Furthermore, we also show that our method can generalize to other types of video data, such as pixel sprites of 2D cartoon characters. Our work presents a new perspective on self-supervised learning of disentangled video representations, contributing to the broader field of video analysis and generation.

视频解耦扩散模型自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。