arXiv:2606.25578cs.CV2026-06中稿 · ECCV

解决大姿态差异下的发型迁移难题,生成更逼真的发型转换结果。

H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks

论文配图:H-Adapter: Pose-Robust Hairstyle Transfer via Attention-Derived, Source-Aligned Hair Masks
图 1 · 摘自论文原文
  • 通过区域损失分离头发与非头发目标,实现空间解耦注意力。
  • 在姿态差异下仍保持最优的FID、CLIP-I等指标,细节还原更精准。
  • 支持文本控制、颜色调节,适合虚拟试穿等实际应用。

发型迁移在虚拟试穿等场景中有实际应用价值,但在源图像与参考图像存在较大头姿差异时仍具挑战性。本文提出H-Adapter,通过区域特定损失训练,分离头发与非头发目标,诱导空间解耦的交叉注意力,进而生成源对齐的发型编辑掩码,用于指导基于扩散模型的修复。在无姿态依赖和姿态差异子集上的实验表明,该方法在姿态差异下实现了最优的FID、FID_CLIP和CLIP-I,同时保持了良好的非头发内容保真度,并显著提升了对精细发型细节的还原质量。此外,H-Adapter还可扩展至文本到图像生成、基于提示的发色控制,兼容身份保留的IP-Adapter变体。我们还引入视觉语言模型作为评判者(VLM-as-a-judge)协议,持续提升发型忠实度、非头发保真度与伪影质量。

原文摘要 · Abstract (English)

Hairstyle transfer has practical applications such as virtual try-on, yet remains challenging when the source and reference exhibit large head-pose discrepancies. We propose H-Adapter, which improves pose robustness by training with a region-specific loss that disentangles hair and non-hair objectives and thereby induces spatially disentangled cross-attention, from which a source-aligned hair edit mask is derived to guide diffusion-based inpainting. Experiments on pose-agnostic and pose-different subsets demonstrate strong quantitative results, including the best FID, $\mathrm{FID}_{\mathrm{CLIP}}$, and CLIP-I under pose differences, while maintaining competitive non-hair preservation and improving qualitative fidelity to fine-grained reference hairstyle details. Beyond source-conditioned transfer, H-Adapter supports practical extensions including text-to-image generation, auxiliary prompt-based hair color control, and compatibility with an identity-preserving IP-Adapter variant. We also introduce a VLM-as-a-judge protocol and observe consistent gains in hairstyle faithfulness, non-hair preservation, and artifact quality.

发型迁移扩散模型姿态鲁棒虚拟试穿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。