arXiv:2511.15092cs.CV2025-11

用多视角信息生成更一致的人体图像,解决单图纹理缺失问题。

Jointly Conditioned Diffusion Model for Multi-View Pose-Guided Person Image Synthesis

  • 通过多视角先验模块和联合条件注入机制融合视图信息。
  • 在多个参考视图下实现更高保真度与跨视角一致性。
  • 适配主流扩散模型,支持任意数量参考图,易部署。

姿态引导的人体图像生成受限于单个参考视图的纹理不完整以及缺乏显式的跨视图交互。本文提出联合条件扩散模型(JCDM),一种利用多视角先验的联合条件扩散框架。外观先验模块(APM)从不完整的参考图像中推断出保持身份一致性的全局先验;联合条件注入(JCI)机制融合多视图线索,并将共享条件注入去噪主干网络,以对齐身份、颜色和纹理。JCDM 支持可变数量的参考视图,仅需少量且有针对性的架构修改即可集成至标准扩散主干网络。实验表明,该方法在保真度与跨视角一致性上达到当前最优水平。

原文摘要 · Abstract (English)

Pose-guided human image generation is limited by incomplete textures from single reference views and the absence of explicit cross-view interaction. We present jointly conditioned diffusion model (JCDM), a jointly conditioned diffusion framework that exploits multi-view priors. The appearance prior module (APM) infers a holistic identity preserving prior from incomplete references, and the joint conditional injection (JCI) mechanism fuses multi-view cues and injects shared conditioning into the denoising backbone to align identity, color, and texture across poses. JCDM supports a variable number of reference views and integrates with standard diffusion backbones with minimal and targeted architectural modifications. Experiments demonstrate state of the art fidelity and cross-view consistency.

图像生成扩散模型多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。