arXiv:2501.06187cs.CV2025-01CVPR被引 66

无需逐个优化,一次生成多个主体的个性化视频。

Multi-subject Open-set Personalization in Video Generation

  • 用交叉注意力融合参考图与文本提示,实现多主体个性化。
  • 在新场景下仍保持高主体保真度,显著优于现有方法。
  • 适合需要快速生成多人物/物体视频的研究者与创作者。

视频个性化技术可生成包含特定人物、宠物或地点的视频,但现有方法通常局限于单一领域、需耗时测试优化,且仅支持单个主体。本文提出 Video Alchemist——一种具备内置多主体、开放集个性化的视频模型,能同时处理前景与背景,无需测试时优化。该模型基于新型 Diffusion Transformer 模块,通过交叉注意力融合条件参考图像及其对应的主题文本提示。构建此类大模型面临两大挑战:数据与评估。由于成对参考图像与视频难以获取,我们采样视频帧作为参考图像,并合成目标视频片段;然而,模型在训练视频中去噪表现良好,却难以泛化至新场景。为此,我们设计了一种自动数据构建流程,结合大量图像增强。其次,开放集视频个性化评估本身极具挑战,因此我们提出了一个聚焦主体保真度的个性化基准,支持多样化个性化场景。大量实验表明,该方法在定量与定性评价上均显著优于现有方法。

原文摘要 · Abstract (English)

Video personalization methods allow us to synthesize videos with specific concepts such as people, pets, and places. However, existing methods often focus on limited domains, require time-consuming optimization per subject, or support only a single subject. We present Video Alchemist $-$ a video model with built-in multi-subject, open-set personalization capabilities for both foreground objects and background, eliminating the need for time-consuming test-time optimization. Our model is built on a new Diffusion Transformer module that fuses each conditional reference image and its corresponding subject-level text prompt with cross-attention layers. Developing such a large model presents two main challenges: dataset and evaluation. First, as paired datasets of reference images and videos are extremely hard to collect, we sample selected video frames as reference images and synthesize a clip of the target video. However, while models can easily denoise training videos given reference frames, they fail to generalize to new contexts. To mitigate this issue, we design a new automatic data construction pipeline with extensive image augmentations. Second, evaluating open-set video personalization is a challenge in itself. To address this, we introduce a personalization benchmark that focuses on accurate subject fidelity and supports diverse personalization scenarios. Finally, our extensive experiments show that our method significantly outperforms existing personalization methods in both quantitative and qualitative evaluations.

视频生成个性化扩散模型多主体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。