解决视频生成中人物身份错乱问题,让真人形象更稳定一致。
Vera: Identity-Faithful Human Subject-to-Video Generation

- 用百万对齐数据训练,精准标注人物身份
- 新方法使人物身份跨帧不变,多人群体不混淆
- 适合需要高保真人像的视频生成场景
主体到视频(S2V)生成在多类别下已取得显著进展,但在以人为主体的生成中,身份一致性仍不足。视频整体可能看起来连贯,但关键人物细节在不同帧、姿态和互动中仍会漂移。这一问题在多人场景中尤为严重,错误的身份角色绑定导致主体混淆、属性互换及过度复制参考图像特征。我们提出Vera,一种统一的人类中心S2V框架,适用于单人与多人生成。首先通过人物级跨片段检索构建百万对齐的人像-视频数据集,提供明确的身份对应关系与多样参考。基于此,Vera引入两项互补设计:身份聚焦掩码监督(IFMS),以空间聚焦方式强化身份感知学习,同时减少无关内容干扰;参考感知层间注意力(RALA),调节视频令牌在DiT主干中与参考身份线索的交互,保持稳定身份锚点并增强分层身份读取能力。大量实验表明,Vera提升了人物身份一致性、多人主体绑定能力与运动自然性,同时降低了身份混淆与过度复制参考图像的现象。
原文摘要 · Abstract (English)
Subject-to-video (S2V) generation has made substantial progress in preserving reference subjects across diverse categories, yet generic subject consistency remains insufficient for human-centric generation. A video may appear globally consistent while identity-critical human details still drift across frames, poses, and interactions. This issue becomes more severe in multi-person scenarios, where incorrect identity-role binding leads to subject confusion, attribute swapping, and excessive copying of reference-specific appearance cues. We propose Vera, a unified human-centric S2V framework for single- and multi-person generation. We first construct a million-pair identity-aligned human image-video dataset through person-level cross-clip retrieval, providing explicit identity correspondence and diverse references. Built on this dataset, Vera introduces two complementary designs. Identity-Focal Masked Supervision (IFMS) strengthens identity-aware learning with spatially focused supervision while reducing interference from irrelevant artifacts. Reference-Aware Layer-wise Attention (RALA) regulates how video tokens interact with reference identity cues in the DiT backbone, preserving stable identity anchors and enhancing layer-aware identity readout. Extensive experiments demonstrate that Vera improves human identity consistency, multi-person subject binding, and motion naturalness, while reducing identity confusion and excessive reference-image copying.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。