arXiv:2603.14770cs.CV2026-03

让多个人脸在指定位置精准生成,且不出现复制粘贴问题。

AnyPhoto: Multi-Person Identity Preserving Image Generation with ID Adaptive Modulation on Location Canvas

  • 用位置对齐的令牌剪枝和旋转编码定位,实现空间精准控制。
  • 通过身份自适应调制,使多人脸保持稳定身份特征。
  • 适合需要多角色精准生成的应用,如虚拟人设计、影视特效。

多人物身份保持生成需在文本提示下将多个参考人脸绑定至指定位置。强身份与布局约束常引发复制粘贴捷径,削弱提示控制能力。我们提出 AnyPhoto,一种基于扩散-变换器微调的框架,包含:(i) RoPE 对齐的位置画布与位置对齐令牌剪枝,实现空间定位;(ii) 基于人脸识别嵌入的 AdaLN 式身份自适应调制,确保身份持续注入;(iii) 身份隔离注意力机制,防止跨身份干扰。训练结合条件流匹配与嵌入空间人脸相似性损失,并引入参考人脸替换与位置画布退化,以抑制捷径。在 MultiID-Bench 上,AnyPhoto 提升了身份相似度,同时降低复制粘贴倾向,且身份数量越多,增益越明显。该方法还支持提示驱动的风格化与精准布局,具备广泛应用潜力。

原文摘要 · Abstract (English)

Multi-person identity-preserving generation requires binding multiple reference faces to specified locations under a text prompt. Strong identity/layout conditions often trigger copy-paste shortcuts and weaken prompt-driven controllability. We present AnyPhoto, a diffusion-transformer finetuning framework with (i) a RoPE-aligned location canvas plus location-aligned token pruning for spatial grounding, (ii) AdaLN-style identity-adaptive modulation from face-recognition embeddings for persistent identity injection, and (iii) identity-isolated attention to prevent cross-identity interference. Training combines conditional flow matching with an embedding-space face similarity loss, together with reference-face replacement and location-canvas degradations to discourage shortcuts. On MultiID-Bench, AnyPhoto improves identity similarity while reducing copy-paste tendency, with gains increasing as the number of identities grows. AnyPhoto also supports prompt-driven stylization with accurate placement, showing great potential application value.

图像生成身份保持扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。