arXiv:2507.05678cs.CV2025-07ICCV被引 7

提出LiON-LoRA,实现视频生成中相机与物体运动的精准协同控制。

LiON-LoRA: Rethinking LoRA Fusion to Unify Controllable Spatial and Temporal Generation for Video Diffusion

  • 通过线性可扩展、正交性与归一化一致性重构LoRA融合机制。
  • 在少量数据下实现轨迹控制精度与运动强度调节的显著提升。
  • 适合需要精细空间与时间运动控制的视频生成研究者。

视频扩散模型(VDMs)通过大规模数据学习,展现出生成逼真视频的强大能力。尽管原始低秩适配(LoRA)可在有限数据下学习特定空间或时间运动以驱动VDM,但实现相机轨迹与物体运动的精确控制仍面临融合不稳定和非线性可扩展性问题。为此,我们提出LiON-LoRA,一个基于线性可扩展性、正交性和归一化一致性三个核心原则的新框架。首先,分析浅层VDM中LoRA特征的正交性,实现低层运动解耦控制;其次,跨层施加归一化一致性以稳定复杂相机运动组合下的融合;第三,将可控标记引入扩散Transformer(DiT),结合改进的自注意力机制,实现对相机与物体运动幅度的线性调控。此外,通过利用静态相机视频,将LiON-LoRA扩展至时间生成,统一了空间与时间控制能力。实验表明,该方法在轨迹控制精度与运动强度调节上优于现有最优方法,且在极小训练数据下仍具优异泛化性。

原文摘要 · Abstract (English)

Video Diffusion Models (VDMs) have demonstrated remarkable capabilities in synthesizing realistic videos by learning from large-scale data. Although vanilla Low-Rank Adaptation (LoRA) can learn specific spatial or temporal movement to driven VDMs with constrained data, achieving precise control over both camera trajectories and object motion remains challenging due to the unstable fusion and non-linear scalability. To address these issues, we propose LiON-LoRA, a novel framework that rethinks LoRA fusion through three core principles: Linear scalability, Orthogonality, and Norm consistency. First, we analyze the orthogonality of LoRA features in shallow VDM layers, enabling decoupled low-level controllability. Second, norm consistency is enforced across layers to stabilize fusion during complex camera motion combinations. Third, a controllable token is integrated into the diffusion transformer (DiT) to linearly adjust motion amplitudes for both cameras and objects with a modified self-attention mechanism to ensure decoupled control. Additionally, we extend LiON-LoRA to temporal generation by leveraging static-camera videos, unifying spatial and temporal controllability. Experiments demonstrate that LiON-LoRA outperforms state-of-the-art methods in trajectory control accuracy and motion strength adjustment, achieving superior generalization with minimal training data. Project Page: https://fuchengsu.github.io/lionlora.github.io/

视频生成扩散模型运动控制LoRA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。