arXiv:2507.13311cs.CV2025-07

用自然语言统一控制服装生成的姿势与光照,实现更真实个性化的虚拟试衣。

FashionPose: Unified Text-Driven Fashion Synthesis with Joint Geometric and Photometric Control

  • 通过双向对比对齐将文字语义映射到几何空间,实现无模板姿态生成。
  • 在4万+图文对数据集上,姿态准确率和真实感显著优于现有方法。
  • 适合需要个性化虚拟试衣、场景自适应生成的电商与设计场景。

真实的可控服装生成对时尚电商至关重要,但需精确协调人体姿态几何与环境光照。传统姿态引导框架存在两大局限:依赖现成估计器的预设骨骼,限制语义灵活性;且多聚焦于中性光照下的影棚生成,难以融合自然语言描述的复杂场景光照。为此,我们提出FashionPose,一种级联架构,在统一的语言驱动界面中实现几何与光照的协同控制。该框架采用解耦但协同的策略:(1) 双向对比对齐机制,将文本语义锚定至显式几何流形,实现无模板姿态生成;(2) 身份锚定合成模块,将几何先验转化为高保真图像并保留细粒度外观;(3) 提示条件重照明模块,以生成姿态为空间锚点,实现环境感知着色。该分层设计有效将高层指令转化为一致视觉表征,确保结构精准与氛围和谐。为支持此范式,我们构建了包含超过40,000个图文-关键点对的PoseCap数据集。大量实验表明,FashionPose在姿态准确性与物理真实性上均超越现有基准,为个性化、场景感知的虚拟时装展示提供稳健解决方案。

原文摘要 · Abstract (English)

Realistic and controllable garment synthesis is essential for fashion e-commerce, yet it demands precise coordination between human pose geometry and environmental photometry. Conventional pose-guided frameworks suffer from two fundamental limitations: they rely heavily on predefined skeletons from off-the-shelf estimators, restricting semantic flexibility; and they predominantly focus on studio-like generation under neutral lighting, failing to reconcile geometric configurations with complex, scene-specific illumination described in natural language. To bridge this gap, we propose FashionPose, a cascaded architecture that reconciles geometric and photometric control within a unified language-driven interface. Unlike conventional frameworks, our framework employs a decoupled yet synergistic strategy: (1) a bidirectional contrastive alignment mechanism that grounds textual semantics into an explicit geometric manifold, enabling template-free pose generation; (2) an identity-anchored synthesis module that translates these geometric priors into high-fidelity imagery while preserving fine-grained appearance; and (3) a prompt-conditioned relighting module that leverages the generated pose as a spatial anchor to achieve environment-aware shading. This hierarchical design effectively transforms high-level instructions into consistent visual representations, ensuring both structural precision and atmospheric harmony. To facilitate this paradigm, we construct PoseCap, a dataset with over 40,000 caption-keypoint pairs. Extensive experiments demonstrate that FashionPose outperforms existing benchmarks in pose accuracy and physical realism, providing a robust solution for personalized, scene-aware virtual fashion displays.

虚拟试衣文本生成光照控制服装合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。