用参考视频当提示词,实现通用可控的视频生成。
Video-As-Prompt: Unified Semantic Control for Video Generation
- 以参考视频为语义提示,通过可插拔专家网络引导冻结模型生成。
- 在10万+数据集上零样本生成,用户偏好率达38.7%,媲美商用模型。
- 无需微调,支持多种任务,适合通用视频生成研究者使用。
视频生成中的统一、可泛化的语义控制仍是重大挑战。现有方法或因结构控制引入不恰当像素级先验而产生伪影,或依赖非泛化的特定条件微调及专用架构。本文提出视频即提示(Video-As-Prompt, VAP),将问题重构为上下文生成。VAP利用参考视频作为直接语义提示,通过即插即用的混合变换器(MoT)专家网络引导冻结的视频扩散变换器(DiT)。该架构避免灾难性遗忘,并采用时间偏置位置编码,消除冗余映射先验,实现稳健上下文检索。为推动此方法并激发未来研究,我们构建了规模最大的语义控制视频生成数据集VAP-Data,包含超过10万对跨100种语义条件的视频对。作为单一统一模型,VAP在开源方法中达到新基准,用户偏好率高达38.7%,媲美领先商用模型。其出色的零样本泛化能力与多下游应用支持,标志着向通用可控视频生成迈出关键一步。
原文摘要 · Abstract (English)
Unified, generalizable semantic control in video generation remains a critical open challenge. Existing methods either introduce artifacts by enforcing inappropriate pixel-wise priors from structure-based controls, or rely on non-generalizable, condition-specific finetuning or task-specific architectures. We introduce Video-As-Prompt (VAP), a new paradigm that reframes this problem as in-context generation. VAP leverages a reference video as a direct semantic prompt, guiding a frozen Video Diffusion Transformer (DiT) via a plug-and-play Mixture-of-Transformers (MoT) expert. This architecture prevents catastrophic forgetting and is guided by a temporally biased position embedding that eliminates spurious mapping priors for robust context retrieval. To power this approach and catalyze future research, we built VAP-Data, the largest dataset for semantic-controlled video generation with over 100K paired videos across 100 semantic conditions. As a single unified model, VAP sets a new state-of-the-art for open-source methods, achieving a 38.7% user preference rate that rivals leading condition-specific commercial models. VAP's strong zero-shot generalization and support for various downstream applications mark a significant advance toward general-purpose, controllable video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。