arXiv:2607.02220cs.CV2026-07

让AI精准生成服装细节图,支持自由聚焦任意部位。

DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation

论文配图:DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation
图 1 · 摘自论文原文
  • 用跨模态特征对齐蒸馏,让模型理解焦点位置与细节对应关系。
  • 在4万+真实配对数据上超越现有开源方法,人评质量显著提升。
  • 适合电商细节生成、虚拟试衣等需高精度局部图像的场景。

基于扩散模型的生成式AI在虚拟试穿、海报生成和产品背景合成等电商业务中取得显著进展。然而,在在线选购服装时,消费者还希望自由查看特定细节区域,如领口、袖口和面料纹理,现有方法未专门研究此需求。为此,我们提出新任务:以焦点条件驱动的服装细节生成,并发布FDBench基准,包含41类、4万+经人工验证的参考-细节配对数据。该任务面临独特语义鸿沟挑战:模型需在无精确提示下,将参考图中的焦点标记与逼真的局部视图对齐,同时保持服装身份一致性。为此,我们提出跨模态特征对齐蒸馏(CFAD),利用微调的DINOv3教师模型,通过双分支蒸馏使多模态扩散变压器在共享语义空间中对齐。为进一步提升生成细节与参考图的一致性,引入一致性奖励模型,联合评估图像对在三个质量维度上的表现,并通过强化学习优化生成过程。实验表明,我们的模型DetailAnywhere在所有指标和人类评估中均显著优于当前最先进的开源方法。

原文摘要 · Abstract (English)

Diffusion-based generative AI has achieved remarkable success in e-commerce applications such as virtual try-on, poster generation, and product background synthesis. However, when making online purchasing decisions for apparel, consumers also desire the freedom to examine specific detail regions of interest, such as collars, cuffs, and fabric textures, yet existing methods have not explicitly studied this setting. We therefore formalize a new, non-template task: Fashion Detail Generation with focus conditioning, and release FDBench, the first benchmark comprising 40K+ human-verified reference-detail pairs across 41 different categories. This task poses a unique semantic gap challenge: the model must bridge the correspondence between a focus marker on a product reference image and a photorealistic close-up view of the indicated region, while faithfully preserving the garment's identity, without any precise prompt. To bridge this gap, we propose Cross-modal Feature Alignment Distillation (CFAD), which leverages a fine-tuned DINOv3 teacher to align both branches of a Multimodal Diffusion Transformer in a shared semantic space via dual-branch distillation. To further improve consistency between generated details and reference images, we introduce a consistency reward model that jointly scores image pairs along three quality axes and optimizes generation via reinforcement learning. Experiments show that our model DetailAnywhere significantly outperforms all state-of-the-art opensource methods across all metrics and human evaluations.

服装生成扩散模型细节生成跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。