arXiv:2503.23951cs.CV2025-03被引 5

同时优化视频外观与运动,避免干扰和背景污染。

JointTuner: Appearance-Motion Adaptive Joint Training for Customized Video Generation

  • 用门控低秩适配动态调节外观或运动学习。
  • 新损失函数提升运动特征,抑制外观干扰。
  • 适配多种模型,适合高质量定制视频生成。

近期定制化视频生成技术在同时适应外观与运动方面取得显著进展。传统方法通常将外观与运动训练解耦,导致概念干扰,出现外观特征或运动模式渲染不准的问题。此外,这些方法常因参考视频中的前景与背景元素混合,引发外观污染。本文提出JointTuner,通过联合优化外观与运动组件缓解上述问题。核心创新包括:门控低秩适配(GLoRA),利用上下文感知激活层动态引导LoRA模块学习外观或运动,保持时空一致性;以及外观无关时间损失(AiT Loss),基于通道-时间移位噪声可抑制外观低频、增强运动高频的特性,使微调时扩散模型更关注运动模式。JointTuner架构无感设计,支持UNet(如ZeroScope)与扩散变压器(如CogVideoX)骨干网络,可随基础视频模型演进而扩展。我们还构建了系统性评估框架,涵盖90种组合,在语义对齐、运动动态性、时间一致性与感知质量四个维度进行评测。

原文摘要 · Abstract (English)

Recent advancements in customized video generation have led to significant improvements in the simultaneous adaptation of appearance and motion. Typically, decoupling the appearance and motion training, prior methods often introduce concept interference, resulting in inaccurate rendering of appearance features or motion patterns. In addition, these methods often suffer from appearance contamination, in which background and foreground elements from reference videos distort the customized video. This paper aims to alleviate these issues by proposing JointTuner. The core motivation of our JointTuner is to enable joint optimization of both appearance and motion components, upon which two key innovations are developed, i.e., Gated Low-Rank Adaptation (GLoRA) and Appearance-independent Temporal Loss (AiT Loss). Specifically, GLoRA uses a context-aware activation layer, analogous to a gating regulator, to dynamically steer LoRA modules toward learning either appearance or motion while maintaining spatio-temporal consistency. Moreover, with the finding that channel-temporal shift noise suppresses appearance-related low-frequencies while enhancing motion-related high-frequencies, we designed the AiT Loss. This loss adds the same shift to the diffusion model's predicted noise during fine-tuning, forcing the model to prioritize learning motion patterns. JointTuner's architecture-agnostic design supports both UNet (e.g., ZeroScope) and Diffusion Transformer (e.g., CogVideoX) backbones, ensuring its customization capabilities scale with the evolution of foundational video models. Furthermore, we present a systematic evaluation framework for appearance-motion combined customization, covering 90 combinations evaluated along four critical dimensions: semantic alignment, motion dynamism, temporal consistency, and perceptual quality. Our project homepage is available online.

视频生成扩散模型联合训练定制化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。