用单图生成高保真个性化视频,人脸相似度更高。
Lynx: Towards High-Fidelity Personalized Video Generation
- 用轻量适配器融合面部嵌入与参考特征,保持身份一致。
- 在40人20个提示的800组测试中,人脸相似度领先。
- 适合需要精准人物还原的视频生成应用。
我们提出Lynx,一种基于开源Diffusion Transformer(DiT)基础模型的高保真个性化视频生成方法。通过引入两个轻量级适配器实现身份一致性:ID-adapter利用Perceiver Resampler将ArcFace提取的面部嵌入转换为紧凑的身份标记进行条件控制;Ref-adapter则整合冻结参考路径中的密集VAE特征,通过跨注意力机制在所有Transformer层注入细粒度细节。该设计在保证时间连贯性和视觉真实性的前提下,显著提升身份保留能力。在包含40名受试者和20个无偏提示的定制基准上进行评估,共产生800个测试案例,结果表明Lynx在人脸相似度、提示遵循能力及视频质量方面均表现优异,推动了个性化视频生成的技术进展。
原文摘要 · Abstract (English)
We present Lynx, a high-fidelity model for personalized video synthesis from a single input image. Built on an open-source Diffusion Transformer (DiT) foundation model, Lynx introduces two lightweight adapters to ensure identity fidelity. The ID-adapter employs a Perceiver Resampler to convert ArcFace-derived facial embeddings into compact identity tokens for conditioning, while the Ref-adapter integrates dense VAE features from a frozen reference pathway, injecting fine-grained details across all transformer layers through cross-attention. These modules collectively enable robust identity preservation while maintaining temporal coherence and visual realism. Through evaluation on a curated benchmark of 40 subjects and 20 unbiased prompts, which yielded 800 test cases, Lynx has demonstrated superior face resemblance, competitive prompt following, and strong video quality, thereby advancing the state of personalized video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。