用参考视频控制生成,实现语义对齐的高效视频合成。
Video2LoRA: Unified Semantic-Controlled Video Generation via Per-Reference-Video LoRA
- 通过轻量超网络为每条语义输入生成个性化LoRA权重。
- 无需微调即可在多种条件下生成语义一致的视频,模型仅150MB。
- 适合需要快速部署、零样本泛化的视频生成场景。
实现多样化视频生成条件间的语义对齐仍是重大挑战。依赖显式结构引导的方法常施加严格空间约束,限制语义灵活性;而针对特定控制类型设计的模型则缺乏互操作性与适应性。为此,我们提出Video2LoRA,一种可扩展、通用的语义控制视频生成框架,以参考视频为条件。该方法利用轻量级超网络为每个语义输入预测个性化LoRA权重,结合辅助矩阵形成自适应LoRA模块,并集成到冻结的扩散主干网络中。这一设计使模型能生成与参考视频语义一致的视频,同时保留关键风格与内容差异,无需任何特定条件训练。最终模型权重小于150MB,存储与部署效率高。Video2LoRA在多种条件下实现连贯、语义对齐的生成,并展现出强大的零样本泛化能力。
原文摘要 · Abstract (English)
Achieving semantic alignment across diverse video generation conditions remains a significant challenge. Methods that rely on explicit structural guidance often enforce rigid spatial constraints that limit semantic flexibility, whereas models tailored for individual control types lack interoperability and adaptability. These design bottlenecks hinder progress toward flexible and efficient semantic video generation. To address this, we propose Video2LoRA, a scalable and generalizable framework for semantic-controlled video generation that conditions on a reference video. Video2LoRA employs a lightweight hypernetwork to predict personalized LoRA weights for each semantic input, which are combined with auxiliary matrices to form adaptive LoRA modules integrated into a frozen diffusion backbone. This design enables the model to generate videos consistent with the reference semantics while preserving key style and content variations, eliminating the need for any per-condition training. Notably, the final model weights less than 150MB, making it highly efficient for storage and deployment. Video2LoRA achieves coherent, semantically aligned generation across diverse conditions and exhibits strong zero-shot generalization to unseen semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。