arXiv:2510.05660cs.CV2025-10ICCV被引 2

无需训练即可将人像无缝插入任意场景,保持身份与细节一致。

Teleportraits: Training-Free People Insertion into Any Scene

  • 利用扩散模型的内在知识,通过逆向生成与无分类器引导实现全局编辑。
  • 仅需单张参考图,就能在复杂场景中精准定位并保留人物身份特征。
  • 首次实现零训练下的人像插入,适合影视合成与虚拟试衣等应用。

将参考图像中的人物真实地插入背景场景是一项极具挑战的任务,需要模型(1)准确判断人物的位置与姿态,(2)根据背景进行高质量个性化。以往方法通常将两者分开处理,且依赖特定训练以达到高性能。本文提出一种统一的零训练流程,基于预训练的文本到图像扩散模型。我们发现扩散模型本身具备在复杂场景中放置人物的知识,无需任务专属训练。通过结合反演技术与无分类器引导,我们的方法实现了感知能力驱动的全局编辑,可无缝将人物插入场景。此外,提出的掩码引导自注意力机制确保了高质量个性化,仅凭单张参考图即可保留人物的身份、服饰和身体特征。据我们所知,这是首个以零训练方式实现真实人像插入的方法,在多样化合成场景图像中达到业界最优水平,同时在背景与主体间保持优异的身份一致性。

原文摘要 · Abstract (English)

The task of realistically inserting a human from a reference image into a background scene is highly challenging, requiring the model to (1) determine the correct location and poses of the person and (2) perform high-quality personalization conditioned on the background. Previous approaches often treat them as separate problems, overlooking their interconnections, and typically rely on training to achieve high performance. In this work, we introduce a unified training-free pipeline that leverages pre-trained text-to-image diffusion models. We show that diffusion models inherently possess the knowledge to place people in complex scenes without requiring task-specific training. By combining inversion techniques with classifier-free guidance, our method achieves affordance-aware global editing, seamlessly inserting people into scenes. Furthermore, our proposed mask-guided self-attention mechanism ensures high-quality personalization, preserving the subject's identity, clothing, and body features from just a single reference image. To the best of our knowledge, we are the first to perform realistic human insertions into scenes in a training-free manner and achieve state-of-the-art results in diverse composite scene images with excellent identity preservation in backgrounds and subjects.

图像合成扩散模型零训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。