通过注意力流场学习,精准控制人物图像生成细节。
Learning Flow Fields in Attention for Controllable Person Image Generation
- 在扩散模型中引入注意力流场正则化,引导查询准确匹配参考图区域。
- 显著降低纹理细节失真,虚拟试穿与姿态迁移效果均达顶尖水平。
- 方法通用性强,可提升其他扩散模型性能,适合可控图像生成研究者。
可控人物图像生成旨在根据参考图像生成目标人物图像,实现外观或姿态的精确控制。然而,现有方法虽整体图像质量高,却常导致参考图中的细粒度纹理细节失真。我们归因于对参考图对应区域的关注不足。为此,提出在注意力中学习流场(Leffa),在训练时显式引导目标查询关注正确的参考键。具体通过在基于扩散模型的基线之上添加注意力图的正则化损失实现。大量实验证明,Leffa在外观控制(虚拟试穿)和姿态迁移(姿态转移)上达到当前最优性能,显著减少细粒度细节失真,同时保持高图像质量。此外,该损失具有模型无关性,可提升其他扩散模型的表现。
原文摘要 · Abstract (English)
Controllable person image generation aims to generate a person image conditioned on reference images, allowing precise control over the person's appearance or pose. However, prior methods often distort fine-grained textural details from the reference image, despite achieving high overall image quality. We attribute these distortions to inadequate attention to corresponding regions in the reference image. To address this, we thereby propose learning flow fields in attention (Leffa), which explicitly guides the target query to attend to the correct reference key in the attention layer during training. Specifically, it is realized via a regularization loss on top of the attention map within a diffusion-based baseline. Our extensive experiments show that Leffa achieves state-of-the-art performance in controlling appearance (virtual try-on) and pose (pose transfer), significantly reducing fine-grained detail distortion while maintaining high image quality. Additionally, we show that our loss is model-agnostic and can be used to improve the performance of other diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。