解决图像生成中身份一致但过度复制的问题,让生成更可控且自然。
WithAnyone: Towards Controllable and ID Consistent Image Generation
- 构建多身份配对数据集MultiID-2M,支持多人场景训练
- 提出对比损失机制,平衡身份保真与姿态表情多样性
- 用户测试验证身份一致性高且生成表达力强
身份一致的图像生成是文本到图像研究的重要方向,现有模型虽能匹配参考身份,但受限于缺乏大规模成对数据集,多数方法依赖重建训练,常导致‘复制粘贴’问题:模型直接复现参考人脸,而非在姿态、表情或光照变化下保持身份一致性。这削弱了可控性并限制生成表现力。为此,我们(1)构建了大规模成对数据集MultiID-2M,专为多人物场景设计;(2)提出新基准,量化复制粘贴现象及身份保真与变化之间的权衡;(3)引入基于对比的身份损失训练范式,利用成对数据平衡保真度与多样性。上述贡献催生WithAnyone——一种扩散模型,在显著减少复制粘贴的同时保持高身份相似性。大量定性与定量实验表明,WithAnyone有效降低复制粘贴现象,提升对姿态和表情的可控性,并维持良好感知质量。用户研究进一步证实其在保持高身份保真度的同时实现富有表现力的可控生成。
原文摘要 · Abstract (English)
Identity-consistent generation has become an important focus in text-to-image research, with recent models achieving notable success in producing images aligned with a reference identity. Yet, the scarcity of large-scale paired datasets containing multiple images of the same individual forces most approaches to adopt reconstruction-based training. This reliance often leads to a failure mode we term copy-paste, where the model directly replicates the reference face rather than preserving identity across natural variations in pose, expression, or lighting. Such over-similarity undermines controllability and limits the expressive power of generation. To address these limitations, we (1) construct a large-scale paired dataset MultiID-2M, tailored for multi-person scenarios, providing diverse references for each identity; (2) introduce a benchmark that quantifies both copy-paste artifacts and the trade-off between identity fidelity and variation; and (3) propose a novel training paradigm with a contrastive identity loss that leverages paired data to balance fidelity with diversity. These contributions culminate in WithAnyone, a diffusion-based model that effectively mitigates copy-paste while preserving high identity similarity. Extensive qualitative and quantitative experiments demonstrate that WithAnyone significantly reduces copy-paste artifacts, improves controllability over pose and expression, and maintains strong perceptual quality. User studies further validate that our method achieves high identity fidelity while enabling expressive controllable generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。