arXiv:2409.09593cs.CV2024-09被引 5

仅用一张图就能稳定实现人像姿态迁移,解决真实场景下的外观一致性问题。

One-Shot Learning for Pose-Guided Person Image Synthesis in the Wild

  • 通过测试时微调预训练文生图模型,仅需单张源图生成结果
  • 引入视觉一致性模块,融合人脸、文本与图像嵌入提升外观统一性
  • 每例处理约48秒,适合实际应用中的快速定制化生成

当前姿态引导人像合成方法严重依赖大量标注三元组数据进行监督训练,但在真实场景中常因训练集与测试样本分布差异而表现不佳。现有方法多通过复杂训练流程、先进架构或更丰富数据集提升泛化能力,本文则采用测试时微调范式,对预训练文生图(Text2Image, T2I)模型进行个性化定制。然而直接微调会导致面部身份与外观属性不一致。为此,我们提出视觉一致性模块(Visual Consistency Module, VCM),通过融合人脸、文本和图像嵌入增强外观一致性。所提方法OnePoseTrans仅需单张源图即可生成高质量姿态迁移结果,在稳定性上优于现有数据驱动方法。每个测试实例在NVIDIA V100 GPU上约需48秒完成模型定制。

原文摘要 · Abstract (English)

Current Pose-Guided Person Image Synthesis (PGPIS) methods depend heavily on large amounts of labeled triplet data to train the generator in a supervised manner. However, they often falter when applied to in-the-wild samples, primarily due to the distribution gap between the training datasets and real-world test samples. While some researchers aim to enhance model generalizability through sophisticated training procedures, advanced architectures, or by creating more diverse datasets, we adopt the test-time fine-tuning paradigm to customize a pre-trained Text2Image (T2I) model. However, naively applying test-time tuning results in inconsistencies in facial identities and appearance attributes. To address this, we introduce a Visual Consistency Module (VCM), which enhances appearance consistency by combining the face, text, and image embedding. Our approach, named OnePoseTrans, requires only a single source image to generate high-quality pose transfer results, offering greater stability than state-of-the-art data-driven methods. For each test case, OnePoseTrans customizes a model in around 48 seconds with an NVIDIA V100 GPU.

姿态迁移文生图零样本学习一致性建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。