无需微调,用重复物体自动构建大规模训练数据,实现高保真物体插入。
ObjectMate: A Recurrence Prior for Object Insertion and Subject-Driven Generation

- 利用大量未标注图像中重复出现的物体,自动生成多视角配对数据
- 在单图或多参考下,身份保留率与光影真实感均优于现有方法
- 无需测试时调优,适合快速生成复杂场景中的物体插入
本文提出一种无需微调的方法,用于物体插入与主体驱动生成。任务目标是根据多个视角输入,将物体自然地融合进由图像或文本指定的场景中。现有方法难以同时实现:(i) 使物体在新场景中具有逼真的姿态与光照表现,(ii) 保持物体身份一致性。我们假设达成此目标需大规模监督,但人工标注成本过高。关键观察是,大量量产物体在大规模无标签数据集中反复出现,涵盖不同场景、姿态与光照。基于此,我们通过检索同一物体的多样化视图,构建强大的成对训练数据集。该数据集支持训练一个简单的文本到图像扩散模型,直接从物体与场景描述生成合成图像。我们在单参考或多参考设置下,对比了ObjectMate与当前最优方法。实验证明,ObjectMate在身份保留和视觉真实性方面表现更优。不同于多数多参考方法,ObjectMate无需测试时调优。
原文摘要 · Abstract (English)
This paper introduces a tuning-free method for both object insertion and subject-driven generation. The task involves composing an object, given multiple views, into a scene specified by either an image or text. Existing methods struggle to fully meet the task's challenging objectives: (i) seamlessly composing the object into the scene with photorealistic pose and lighting, and (ii) preserving the object's identity. We hypothesize that achieving these goals requires large scale supervision, but manually collecting sufficient data is simply too expensive. The key observation in this paper is that many mass-produced objects recur across multiple images of large unlabeled datasets, in different scenes, poses, and lighting conditions. We use this observation to create massive supervision by retrieving sets of diverse views of the same object. This powerful paired dataset enables us to train a straightforward text-to-image diffusion architecture to map the object and scene descriptions to the composited image. We compare our method, ObjectMate, with state-of-the-art methods for object insertion and subject-driven generation, using a single or multiple references. Empirically, ObjectMate achieves superior identity preservation and more photorealistic composition. Differently from many other multi-reference methods, ObjectMate does not require slow test-time tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。