arXiv:2512.15069cs.CVcs.AI2025-12被引 1

用多视角+姿态图+文本生成一致且可控的人像,解决服装变形和对齐问题。

PMMD: A pose-guided multi-view multi-modal diffusion for person generation

  • 融合多视角图像、姿态图与文本,统一建模人体外观与动作
  • 在DeepFashion数据集上显著提升图像一致性与细节保真度
  • 适合虚拟试衣、数字人生成等需要高可控性的场景

生成具有可控姿态和外观的一致人像,在虚拟试衣、图像编辑和数字人创建中至关重要。现有方法常面临遮挡、服装风格漂移和姿态对齐偏差问题。本文提出姿态引导的多视角多模态扩散模型(PMMD),通过多视图参考、姿态图和文本提示生成逼真人像。设计多模态编码器联合建模视觉视角、姿态特征与语义描述,降低跨模态差异,提升身份保真度;引入ResCVA模块增强局部细节并保持全局结构;设计跨模态融合模块,在去噪过程中整合图像语义与文本信息。在DeepFashion MultiModal数据集上的实验表明,PMMD在一致性、细节保留和可控性方面优于代表性基线方法。项目主页与代码见https://github.com/ZANMANGLOOPYE/PMMD。

原文摘要 · Abstract (English)

Generating consistent human images with controllable pose and appearance is essential for applications in virtual try on, image editing, and digital human creation. Current methods often suffer from occlusions, garment style drift, and pose misalignment. We propose Pose-guided Multi-view Multimodal Diffusion (PMMD), a diffusion framework that synthesizes photorealistic person images conditioned on multi-view references, pose maps, and text prompts. A multimodal encoder jointly models visual views, pose features, and semantic descriptions, which reduces cross modal discrepancy and improves identity fidelity. We further design a ResCVA module to enhance local detail while preserving global structure, and a cross modal fusion module that integrates image semantics with text throughout the denoising pipeline. Experiments on the DeepFashion MultiModal dataset show that PMMD outperforms representative baselines in consistency, detail preservation, and controllability. Project page and code are available at https://github.com/ZANMANGLOOPYE/PMMD.

人像生成扩散模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。