arXiv:2503.08133cs.CVcs.AI2025-03被引 1

用多模态引导提升扩散模型生成真实手部图像质量。

MGHanD: Multi-modal Guidance for authentic Hand Diffusion

  • 引入视觉与文本双重引导,优化手部结构生成
  • 在无需先验条件情况下,显著提升手部细节真实性
  • 适合需要高质量手部图像的生成任务使用

基于扩散的方法在文本到图像生成中取得显著进展,能从文本提示生成逼真图像。然而,这些模型在生成真实手部时仍面临挑战,常出现手指数量错误或结构扭曲问题。MGHanD通过在推理过程中引入多模态引导解决此问题。视觉引导采用在包含配对真实与生成图像及描述的多个野外手部数据集上训练的判别器;文本引导则使用LoRA适配器,在潜在空间学习从‘手’到更详细提示如‘自然手’、‘解剖正确手指’的方向。通过逐步扩大的手部掩码在指定时间步施加引导,实现手部精细化同时保持预训练模型的强大生成能力。实验表明,该方法在无需特定条件或先验的前提下,显著提升手部生成质量。我们进行了定量、定性评估及用户研究,验证了该方法在生成高质量手部图像方面的优势。

原文摘要 · Abstract (English)

Diffusion-based methods have achieved significant successes in T2I generation, providing realistic images from text prompts. Despite their capabilities, these models face persistent challenges in generating realistic human hands, often producing images with incorrect finger counts and structurally deformed hands. MGHanD addresses this challenge by applying multi-modal guidance during the inference process. For visual guidance, we employ a discriminator trained on a dataset comprising paired real and generated images with captions, derived from various hand-in-the-wild datasets. We also employ textual guidance with LoRA adapter, which learns the direction from `hands' towards more detailed prompts such as `natural hands', and `anatomically correct fingers' at the latent level. A cumulative hand mask which is gradually enlarged in the assigned time step is applied to the added guidance, allowing the hand to be refined while maintaining the rich generative capabilities of the pre-trained model. In the experiments, our method achieves superior hand generation qualities, without any specific conditions or priors. We carry out both quantitative and qualitative evaluations, along with user studies, to showcase the benefits of our approach in producing high-quality hand images.

手部生成扩散模型多模态引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。