用流匹配模型实现秒级人物图像实时生成,兼顾速度与质量。
Real-Time Person Image Synthesis Using a Flow Matching Model
- 采用流匹配机制,训练和采样更高效稳定。
- 在DeepFashion上实现近实时生成,速度提升超2倍。
- 适合直播、AR/VR等需即时反馈的交互场景。
姿态引导的人物图像合成(PGPIS)根据目标姿态和源图像生成真实感人物图像,在手语视频生成、AR/VR、游戏和直播等场景中至关重要。然而,高保真图像生成面临复杂人体姿态动态变化带来的实时性挑战。尽管基于扩散模型的方法图像质量优异,但采样速度慢,难以满足实时应用需求,如直播中的手语视频生成。为此,本文提出基于流匹配(Flow Matching)的实时人物图像合成方法(RPFM),支持条件生成并可在潜在空间运行,显著提升训练与采样效率。在广泛使用的DeepFashion数据集上的实验表明,RPFM实现近实时采样速度,生成速度较当前最优模型提升超过2倍,仅以可接受的轻微精度损失换取实时性能,适用于对速度与质量均敏感的实时交互系统。
原文摘要 · Abstract (English)
Pose-Guided Person Image Synthesis (PGPIS) generates realistic person images conditioned on a target pose and a source image. This task plays a key role in various real-world applications, such as sign language video generation, AR/VR, gaming, and live streaming. In these scenarios, real-time PGPIS is critical for providing immediate visual feedback and maintaining user immersion.However, achieving real-time performance remains a significant challenge due to the complexity of synthesizing high-fidelity images from diverse and dynamic human poses. Recent diffusion-based methods have shown impressive image quality in PGPIS, but their slow sampling speeds hinder deployment in time-sensitive applications. This latency is particularly problematic in tasks like generating sign language videos during live broadcasts, where rapid image updates are required. Therefore, developing a fast and reliable PGPIS model is a crucial step toward enabling real-time interactive systems. To address this challenge, we propose a generative model based on flow matching (FM). Our approach enables faster, more stable, and more efficient training and sampling. Furthermore, the proposed model supports conditional generation and can operate in latent space, making it especially suitable for real-time PGPIS applications where both speed and quality are critical. We evaluate our proposed method, Real-Time Person Image Synthesis Using a Flow Matching Model (RPFM), on the widely used DeepFashion dataset for PGPIS tasks. Our results show that RPFM achieves near-real-time sampling speeds while maintaining performance comparable to the state-of-the-art models. Our methodology trades off a slight, acceptable decrease in generated-image accuracy for over a twofold increase in generation speed, thereby ensuring real-time performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。