arXiv:2510.07340cs.GRcs.LG2025-10被引 3

通过分离特征空间干扰,实现高效精准的人物图像生成与编辑。

SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation

  • 利用正交约束在特征空间中分离主体身份,避免无关因素干扰。
  • 仅用1万样本训练,在主体保留和可控编辑上优于已有方法。
  • 适用于需要高保真人物生成的创意设计、虚拟角色开发场景。

个性化图像生成旨在忠实保留参考主体的身份,同时适配多样的文本提示。现有基于优化的方法虽能保证高保真度,但计算成本高昂;而基于学习的方法虽效率更高,却因受噪声因素影响导致表征纠缠。本文提出SpotDiff,一种新型学习型方法,通过在特征空间中识别并解耦干扰,提取主体特异性特征。借助预训练的CLIP图像编码器及专门用于姿态和背景的专家网络,SpotDiff通过特征空间的正交性约束实现主体身份的隔离。为支持可解释训练,我们构建了SpotDiff10k数据集,包含一致的姿态与背景变化。实验表明,SpotDiff在主体保留鲁棒性和可控编辑方面优于先前方法,且仅需10,000个训练样本即达到竞争性性能。

原文摘要 · Abstract (English)

Personalized image generation aims to faithfully preserve a reference subject's identity while adapting to diverse text prompts. Existing optimization-based methods ensure high fidelity but are computationally expensive, while learning-based approaches offer efficiency at the cost of entangled representations influenced by nuisance factors. We introduce SpotDiff, a novel learning-based method that extracts subject-specific features by spotting and disentangling interference. Leveraging a pre-trained CLIP image encoder and specialized expert networks for pose and background, SpotDiff isolates subject identity through orthogonality constraints in the feature space. To enable principled training, we introduce SpotDiff10k, a curated dataset with consistent pose and background variations. Experiments demonstrate that SpotDiff achieves more robust subject preservation and controllable editing than prior methods, while attaining competitive performance with only 10k training samples.

图像生成特征解耦个性化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。