arXiv:2504.00396cs.CV2025-04被引 1

让AI画人像时只改目标特征,不破坏原模型能力。

SPF-Portrait: Towards Pure Text-to-Portrait Customization with Semantic Pollution-Free Fine-Tuning

  • 双路径对比学习,用原模型做行为对齐参考。
  • 新设计的语义控制图精准引导特征响应区域。
  • 适合需要精细定制人像又怕模型变味的开发者。

在定制化人像生成中,现有微调方法常严重改变预训练文本到图像(T2I)模型的原始行为(如身份、构图等)。为解决此问题,我们提出SPF-Portrait,首个实现纯语义无污染微调的方案。其设计双路径对比学习框架,将原模型作为行为对齐参考。引入新型语义感知细控图,指示目标语义响应区域强度,空间化引导对比路径间的对齐过程,自适应平衡各区域的行为一致性与目标语义响应性。此外,提出响应增强机制,强化目标语义表达,缓解跨模态监督带来的表征差异。通过上述策略,实现纯文本到人像定制化的增量语义学习。大量实验表明,SPF-Portrait达到当前最佳性能。

原文摘要 · Abstract (English)

Fine-tuning a pre-trained Text-to-Image (T2I) model on a tailored portrait dataset is the mainstream method for text-to-portrait customization. However, existing methods often severely impact the original model's behavior (e.g., changes in ID, layout, etc.) while customizing portrait attributes. To address this issue, we propose SPF-Portrait, a pioneering work to purely understand customized target semantics and minimize disruption to the original model. In our SPF-Portrait, we design a dual-path contrastive learning pipeline, which introduces the original model as a behavioral alignment reference for the conventional fine-tuning path. During the contrastive learning, we propose a novel Semantic-Aware Fine Control Map that indicates the intensity of response regions of the target semantics, to spatially guide the alignment process between the contrastive paths. It adaptively balances the behavioral alignment across different regions and the responsiveness of the target semantics. Furthermore, we propose a novel response enhancement mechanism to reinforce the presentation of target semantics, while mitigating representation discrepancy inherent in direct cross-modal supervision. Through the above strategies, we achieve incremental learning of customized target semantics for pure text-to-portrait customization. Extensive experiments show that SPF-Portrait achieves state-of-the-art performance. Project page: https://spf-portrait.github.io/SPF-Portrait/

文本生成人像定制微调优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。