提升扩散模型生成人脸手部质量,兼顾全局与局部细节
FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
- 设计多目标微调框架,引入人脸手部位置先验增强局部监督
- 在多个数据集上显著提升手部和面部生成质量,保持整体图像一致
- 适合关注人体生成细节、追求公平优化的研究者与开发者
图像生成随着大规模文本到图像模型的发展取得了显著进展,尤其是基于扩散模型的方法。然而,在训练过程中对局部区域(如人脸、手部)的监督不足,导致生成真实感的人体图像仍具挑战性。为此,我们提出 FairHuman,一种多目标微调方法,旨在公平提升全局与局部生成质量。具体地,构建三个学习目标:一个源自默认扩散目标函数的全局目标,以及基于预标注位置先验的两个局部目标(分别针对手部和面部)。随后,基于最小潜在延迟(MPD)准则推导最优参数更新策略,实现多目标问题下的公平优化。实验表明,该方法在不同场景下显著改善了复杂局部细节的生成效果,同时维持整体图像质量。
原文摘要 · Abstract (English)
Image generation has achieved remarkable progress with the development of large-scale text-to-image models, especially diffusion-based models. However, generating human images with plausible details, such as faces or hands, remains challenging due to insufficient supervision of local regions during training. To address this issue, we propose FairHuman, a multi-objective fine-tuning approach designed to enhance both global and local generation quality fairly. Specifically, we first construct three learning objectives: a global objective derived from the default diffusion objective function and two local objectives for hands and faces based on pre-annotated positional priors. Subsequently, we derive the optimal parameter updating strategy under the guidance of the Minimum Potential Delay (MPD) criterion, thereby attaining fairness-ware optimization for this multi-objective problem. Based on this, our proposed method can achieve significant improvements in generating challenging local details while maintaining overall quality. Extensive experiments showcase the effectiveness of our method in improving the performance of human image generation under different scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。