arXiv:2606.26668cs.CVcs.AI2026-06被引 1

让视频生成同时控制内容、风格和动作,效果更精准可控。

Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization

论文配图:Disco-LoRA: Disentangled Composition of Content, Style, and Motion for Multi-concept Video Customization
图 1 · 摘自论文原文
  • 分两阶段拆解内容、风格、动作的分离与重组。
  • 通过统计正则化使不同LoRA权重协调一致,避免干扰。
  • 适合需要精细控制视频元素的研究者或创作者。

基于文本到视频(T2V)模型的视频定制旨在从参考数据中学习特定特征以生成可控制的视频。尽管图像风格迁移和视频动作定制已取得显著进展,但同时控制内容、风格和动作等多重概念仍是重大挑战。本文系统定义了多概念视频定制任务,需联合控制内容、风格和动作。为此,我们构建了一个全面的基准,并提出Disco-LoRA统一框架,通过两阶段解耦与灵活重组实现该目标:(1) 将目标分解为内容-风格与内容-动作两个子任务,分别采用迭代双LoRA解耦框架有效分离不同概念;(2) 发现层间权重趋势决定LoRA身份,而权重幅值决定可组合性,因此提出基于Z-score的统计正则化,对齐权重分布,在保留层间趋势的同时最小化不同LoRA间的干扰。大量实验表明,Disco-LoRA在多概念视频定制中表现优异,能有效保持外观、风格和动作,实现可控的文本到视频生成。

原文摘要 · Abstract (English)

Video customization based on Text-to-Video (T2V) models aims to learn specific features from reference data to generate controllable videos. While significant strides have been made in image stylization and video motion customization, simultaneously controlling multiple concepts, such as content, style, and motion, remains a major challenge. In this work, we systematically define the task of multi-concept video customization, which requires the joint control of content, style, and motion. To facilitate research in this area, we construct a comprehensive benchmark and propose Disco-LoRA, a unified framework designed to tackle this problem by disentangling and flexibly recombining different concepts in two stages: (1) We decompose the objective into two sub-tasks: Content-Style and Content-Motion. Each sub-task is addressed using our Iterative Dual-LoRA Disentanglement Framework, which effectively disentangles distinct concepts within the data. (2) We identify layer-wise weight trends as crucial for LoRA identity, while weight magnitudes dictate composability. To harmonize these scales, we propose a Z-score-based statistical regularization that aligns weight distributions, preserving layer-wise trends while minimizing interference between different LoRAs. Extensive experiments show that Disco-LoRA excels in multi-concept video customization, effectively preserving appearance, style, and motion for controllable text-to-video generation.

视频生成风格控制LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。