用多视角身份拼图实现人像视频生成的稳定识别。
ARGUS: Stacked Multi-View Identity Mosaic Injection for Subject-Preserving Video Generation

- 将身份信息转为动态拼图注入扩散模型,避免单图依赖。
- 在公开数据集上达64.38分,大角度和遮挡下识别提升超15点。
- 适合需要精准人物保持的视频生成场景,如影视与虚拟人。
仅靠正面人脸相似性无法解决人物保持的视频生成问题:生成人物必须在运动、大视角变化、表情切换、遮挡、尺度变化以及文本、首帧与身份参考之间的冲突中仍可识别。我们指出核心瓶颈在于点引用范式,该范式将身份压缩为单一静态观测,与姿态、配饰、光照、背景和相机统计纠缠。本文提出Argus框架,核心是堆叠多视角身份拼图注入(SMII):将多模态大模型选择的身份证据转换为3×3拼图,同步扩散时间,并以只读记忆形式注入到万能注意力网络(Wan)的原生标记空间。这使身份从外部干净适配器或单张参考图转变为紧凑的动态分布。围绕SMII,MLLM身份导演选择关键身份时刻并解决条件冲突;无跨对反事实训练、时序身份退火与自相似性引导增强鲁棒性,无需成对主体-视频监督。我们还发布了硬身份-名人(HardID-Celeb)基准,并引入俯仰角得分(YawScore)与遮挡得分(OccScore)评估大俯仰角与首帧遮挡下的表现。Argus在OpenS2V-Eval人类领域达到64.38总分、71.86人脸相似度、51.62联合得分与79.14自然度得分。在HardID-Celeb上,人脸相似度达76.80,俯仰角与遮挡得分分别较最强基线提升12.60和15.10点,证明动态身份记忆与大规模反事实自监督对人物保持视频生成高度有效。
原文摘要 · Abstract (English)
Subject-preserving video generation is not solved by frontal-face similarity alone: a generated person must remain recognizable across motion, large viewpoint changes, expression shifts, occlusion, scale variation, and conflicts among text, first-frame, and identity references. We argue that the central bottleneck is the point-reference paradigm, which collapses identity into a single static observation entangled with pose, accessories, lighting, background, and camera statistics. We introduce Argus, a Wan-based framework centered on Stacked Multi-View Identity Mosaic Injection (SMII). SMII converts MLLM-selected image/video identity evidence into a 3*3 stacked mosaic, synchronizes the mosaic with the current diffusion time, and injects it as negative-time read-only memory in Wan's native token space. This turns identity from an external clean adapter or a single reference image into a compact dynamic distribution. Around SMII, an MLLM Identity Director selects informative identity moments and resolves condition conflicts, while no-cross-pair counterfactual training, Temporal Identity Annealing, and Adaptive Self-Likeness Guidance improve robustness without paired subject-video supervision. We further release HardID-Celeb, a public-figure identity-stress benchmark, and introduce YawScore and OccScore to probe large-yaw and first-frame-occlusion robustness. Argus achieves state-of-the-art results on OpenS2V-Eval Human-Domain, reaching 64.38 Total Score, 71.86 FaceSim, 51.62 NexusScore, and 79.14 NaturalScore. On HardID-Celeb, Argus obtains 76.80 FaceSim and improves YawScore and OccScore by 12.60 and 15.10 points over the strongest baselines, demonstrating that dynamic identity memory and large-scale counterfactual self-supervision are highly effective for subject-preserving video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。