arXiv:2512.06905cs.CV2025-12被引 10

无需参考视频对,零样本生成保持身份一致的视频。

Scaling Zero-Shot Reference-to-Video Generation

  • 仅用视频-文本对训练,通过掩码策略学习身份一致表征
  • 在多个参考图像下仍保持高质量生成,超越有监督方法
  • 适合需要快速扩展参考生成能力的研究者

参考到视频(R2V)生成旨在根据文本提示合成与参考图像保持主体身份一致的视频。然而现有方法依赖于昂贵且难以扩展的参考图像-视频-文本三元组数据。本文提出 Saber,一种可扩展的零样本框架,无需显式 R2V 数据。Saber 仅在视频-文本对上训练,采用掩码训练策略和基于注意力的模型设计,学习具有一致身份和参考感知的表征。进一步引入掩码增强技术,缓解参考生成中常见的拼贴伪影。Saber 在不同数量参考图像下均展现优异泛化能力,在 OpenS2V-Eval 基准测试中优于使用 R2V 数据训练的方法。

原文摘要 · Abstract (English)

Reference-to-video (R2V) generation aims to synthesize videos that align with a text prompt while preserving the subject identity from reference images. However, current R2V methods are hindered by the reliance on explicit reference image-video-text triplets, whose construction is highly expensive and difficult to scale. We bypass this bottleneck by introducing Saber, a scalable zero-shot framework that requires no explicit R2V data. Trained exclusively on video-text pairs, Saber employs a masked training strategy and a tailored attention-based model design to learn identity-consistent and reference-aware representations. Mask augmentation techniques are further integrated to mitigate copy-paste artifacts common in reference-to-video generation. Moreover, Saber demonstrates remarkable generalization capabilities across a varying number of references and achieves superior performance on the OpenS2V-Eval benchmark compared to methods trained with R2V data.

视频生成零样本身份保持

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。