无需微调,一键生成个性化动态视频。
Zero-Shot Dynamic Concept Personalization with Grid-Based LoRA
- 用2×2网格结构组织视频对,训练轻量级LoRA适配器。
- 单次前向传播即可生成连贯且身份一致的动态视频。
- 零样本泛化,适用于未见过的角色和动作场景。
文本到视频生成技术已能高质量地根据文本或图像提示生成视频。尽管动态概念个性化(从单个视频中捕捉特定主体的外观与运动)现已可行,但现有方法大多需针对每个实例微调,难以扩展。本文提出一种完全零样本的动态概念个性化框架。方法利用2×2视频网格,空间化组织输入输出对,训练轻量级Grid-LoRA适配器,实现网格内编辑与组合。推理时,专用的Grid Fill模块补全部分观测布局,生成时间连贯且身份保持的输出。模型训练完成后,整个系统仅需一次前向传播,即可在不进行任何测试时优化的情况下,泛化至未见的动态概念。大量实验表明,该方法在多种未训练过的主体和编辑场景中均能生成高质量且一致的结果。
原文摘要 · Abstract (English)
Recent advances in text-to-video generation have enabled high-quality synthesis from text and image prompts. While the personalization of dynamic concepts, which capture subject-specific appearance and motion from a single video, is now feasible, most existing methods require per-instance fine-tuning, limiting scalability. We introduce a fully zero-shot framework for dynamic concept personalization in text-to-video models. Our method leverages structured 2x2 video grids that spatially organize input and output pairs, enabling the training of lightweight Grid-LoRA adapters for editing and composition within these grids. At inference, a dedicated Grid Fill module completes partially observed layouts, producing temporally coherent and identity preserving outputs. Once trained, the entire system operates in a single forward pass, generalizing to previously unseen dynamic concepts without any test-time optimization. Extensive experiments demonstrate high-quality and consistent results across a wide range of subjects beyond trained concepts and editing scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。