解决个性化图像生成中的过拟合与评估偏差问题
Towards More Accurate Personalized Image Generation: Addressing Overfitting and Evaluation Bias
- 引入吸引子机制过滤训练图像干扰,聚焦主体表征学习
- 新基准测试集使自动评估更可靠,避免性能高估
- 适合关注个性化生成模型真实性能的研究者
通过文本提示实现个性化图像生成有望提升日常生活与专业工作中的视觉内容定制效率。个性化图像生成的目标是在保持主体一致性的前提下,响应多样化的文本描述。然而,现有方法在忠实遵循文本提示的同时,容易对训练数据过拟合。本文提出一种新型训练流程,引入吸引子机制过滤训练图像中的干扰信息,使模型更专注学习个性化主体的有效表征。此外,当前评估方法因缺乏专用测试集而存在偏差:通常使用个性化任务的训练数据计算文本-图像和图像-图像相似度,虽有参考价值,但易高估模型表现;人工评估虽为替代方案,却常受主观偏见与不一致影响。为此,我们构建了一个多样化且高质量的测试集,搭配精心设计的提示,形成新基准。在此基础上,自动评估指标可更准确地衡量模型性能。
原文摘要 · Abstract (English)
Personalized image generation via text prompts has great potential to improve daily life and professional work by facilitating the creation of customized visual content. The aim of image personalization is to create images based on a user-provided subject while maintaining both consistency of the subject and flexibility to accommodate various textual descriptions of that subject. However, current methods face challenges in ensuring fidelity to the text prompt while not overfitting to the training data. In this work, we introduce a novel training pipeline that incorporates an attractor to filter out distractions in training images, allowing the model to focus on learning an effective representation of the personalized subject. Moreover, current evaluation methods struggle due to the lack of a dedicated test set. The evaluation set-up typically relies on the training data of the personalization task to compute text-image and image-image similarity scores, which, while useful, tend to overestimate performance. Although human evaluations are commonly used as an alternative, they often suffer from bias and inconsistency. To address these issues, we curate a diverse and high-quality test set with well-designed prompts. With this new benchmark, automatic evaluation metrics can reliably assess model performance
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。