arXiv:2410.23332cs.CVcs.AI2024-10NeurIPS被引 17

用百万级人像数据+低秩专家混合,提升扩散模型生成人脸手部的自然度

MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts

  • 构建百万级人像数据集,细分面部与手部近景图像增强先验
  • 提出低秩专家混合方法,按需调用特定部位优化模块
  • 在人脸手部生成上超越现有方法,适合高精度人像生成场景

文本到图像扩散模型因强大的生成能力受到广泛关注。然而,在涉及人脸和手部的人像生成任务中,由于训练先验不足,生成结果往往不够自然。本文从数据和方法两方面改进:1)收集超过一百万张高质量人像场景图像,以及专门针对面部和手部的近景图像数据集,构建丰富的人像先验知识库;2)提出一种简单有效的低秩专家混合方法(MoLE),将分别在面部和手部近景数据上训练的低秩模块视为专家,依据生成尺度动态调用以优化对应区域。通过构建两个基准测试,结合多种量化指标与人工评估验证了该方法在人像生成上的优越性。相关数据集、模型及代码已公开。

原文摘要 · Abstract (English)

Text-to-image diffusion has attracted vast attention due to its impressive image-generation capabilities. However, when it comes to human-centric text-to-image generation, particularly in the context of faces and hands, the results often fall short of naturalness due to insufficient training priors. We alleviate the issue in this work from two perspectives. 1) From the data aspect, we carefully collect a human-centric dataset comprising over one million high-quality human-in-the-scene images and two specific sets of close-up images of faces and hands. These datasets collectively provide a rich prior knowledge base to enhance the human-centric image generation capabilities of the diffusion model. 2) On the methodological front, we propose a simple yet effective method called Mixture of Low-rank Experts (MoLE) by considering low-rank modules trained on close-up hand and face images respectively as experts. This concept draws inspiration from our observation of low-rank refinement, where a low-rank module trained by a customized close-up dataset has the potential to enhance the corresponding image part when applied at an appropriate scale. To validate the superiority of MoLE in the context of human-centric image generation compared to state-of-the-art, we construct two benchmarks and perform evaluations with diverse metrics and human studies. Datasets, model, and code are released at https://sites.google.com/view/mole4diffuser/.

扩散模型人像生成低秩专家图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。