arXiv:2508.07901cs.CV2025-08被引 23

用极少参数实现视频生成中身份精准控制,还能无缝对接其他功能。

Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation

  • 在预训练模型中加入条件图像分支,通过受限自注意力控制身份。
  • 仅用2000对数据、1%额外参数即超越全参数训练方法。
  • 轻量易插拔,适合需要身份一致性的视频生成任务。

生成与用户指定身份一致的高保真人像视频在生成式AI中至关重要但极具挑战。现有方法通常依赖大量训练参数,且难以与其他AIGC工具兼容。本文提出Stand-In,一种轻量级、可即插即用的身份控制框架。具体地,我们在预训练视频生成模型中引入条件图像分支,通过条件位置映射的受限自注意力实现身份控制。得益于这一设计,模型能有效保留预训练先验,仅需约1%的额外参数和2000对训练数据,便在视频质量与身份一致性上超越全参数微调方法。此外,该框架可无缝集成至主体驱动视频生成、姿态参考视频生成、风格化及人脸替换等任务中。

原文摘要 · Abstract (English)

Generating high-fidelity human videos that match user-specified identities is important yet challenging in the field of generative AI. Existing methods often rely on an excessive number of training parameters and lack compatibility with other AIGC tools. In this paper, we propose Stand-In, a lightweight and plug-and-play framework for identity preservation in video generation. Specifically, we introduce a conditional image branch into the pre-trained video generation model. Identity control is achieved through restricted self-attentions with conditional position mapping. Thanks to these designs, which greatly preserve the pre-trained prior of the video generation model, our approach is able to outperform other full-parameter training methods in video quality and identity preservation, even with just $\sim$1% additional parameters and only 2000 training pairs. Moreover, our framework can be seamlessly integrated for other tasks, such as subject-driven video generation, pose-referenced video generation, stylization, and face swapping.

视频生成身份控制轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。