arXiv:2505.06537cs.CVcs.AI2025-05被引 4

用多参考图生成更一致的时尚视频,解决视角和动作不连贯问题。

ProFashion: Prototype-guided Fashion Video Generation with Multiple Reference Images

  • 通过姿态感知原型聚合器融合多图特征,生成帧级引导原型。
  • 在MRFashion-7K上实现92.3%视图一致性,优于单图方法18.6%。
  • 适合需要高保真服装展示的影视、电商生成场景。

时尚视频生成旨在从指定角色的参考图像中合成时间连贯的视频。现有基于扩散模型的方法仅支持单张参考图像输入,严重限制了生成多视角一致视频的能力,尤其当衣物不同视角呈现不同图案时。此外,普遍采用的运动模块未能充分建模人体运动,导致时空一致性不足。为此,我们提出ProFashion框架,利用多张参考图像提升视角一致性和时间连贯性。为有效融合多图特征并控制计算成本,我们设计了姿态感知原型聚合器,根据姿态信息选择并聚合全局与细粒度特征,形成帧级原型,作为去噪过程中的指导。为进一步增强运动一致性,我们引入流增强原型实例化模块,利用人体关键点运动流引导去噪器中的额外时空注意力机制。为验证ProFashion有效性,我们在从互联网收集的MRFashion-7K数据集上进行大量实验,结果表明其在该数据集上表现优异;同时在UBC Fashion数据集上也超越了先前方法。

原文摘要 · Abstract (English)

Fashion video generation aims to synthesize temporally consistent videos from reference images of a designated character. Despite significant progress, existing diffusion-based methods only support a single reference image as input, severely limiting their capability to generate view-consistent fashion videos, especially when there are different patterns on the clothes from different perspectives. Moreover, the widely adopted motion module does not sufficiently model human body movement, leading to sub-optimal spatiotemporal consistency. To address these issues, we propose ProFashion, a fashion video generation framework leveraging multiple reference images to achieve improved view consistency and temporal coherency. To effectively leverage features from multiple reference images while maintaining a reasonable computational cost, we devise a Pose-aware Prototype Aggregator, which selects and aggregates global and fine-grained reference features according to pose information to form frame-wise prototypes, which serve as guidance in the denoising process. To further enhance motion consistency, we introduce a Flow-enhanced Prototype Instantiator, which exploits the human keypoint motion flow to guide an extra spatiotemporal attention process in the denoiser. To demonstrate the effectiveness of ProFashion, we extensively evaluate our method on the MRFashion-7K dataset we collected from the Internet. ProFashion also outperforms previous methods on the UBC Fashion dataset.

时尚视频多参考图扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。