首帧可当记忆库,用少量样本就能定制视频内容。
First Frame Is the Place to Go for Video Content Customization
- 利用首帧作为视觉信息记忆,实现内容重用
- 仅需20-50个样本即可完成定制,无需模型修改
- 适合需要快速个性化视频生成的场景
视频生成模型中的首帧传统上被视为时空起始点,仅作为后续动画的种子。本文揭示了一个根本性新视角:视频模型隐式地将首帧当作概念记忆缓冲区,用于存储视觉实体以供生成过程中复用。基于此发现,我们证明仅使用20-50个训练样本,在不改变架构或进行大规模微调的情况下,即可在多种场景中实现鲁棒且泛化的视频内容定制。这一结果揭示了视频生成模型在参考式视频定制中被忽视的强大能力。
原文摘要 · Abstract (English)
What role does the first frame play in video generation models? Traditionally, it's viewed as the spatial-temporal starting point of a video, merely a seed for subsequent animation. In this work, we reveal a fundamentally different perspective: video models implicitly treat the first frame as a conceptual memory buffer that stores visual entities for later reuse during generation. Leveraging this insight, we show that it's possible to achieve robust and generalized video content customization in diverse scenarios, using only 20-50 training examples without architectural changes or large-scale finetuning. This unveils a powerful, overlooked capability of video generation models for reference-based video customization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。