arXiv:2511.15197cs.CV2025-11被引 1

零样本生成新框架,让物体无缝融入不同风格场景

Insert In Style: A Zero-Shot Generative Framework for Harmonious Cross-Domain Object Composition

  • 分阶段训练分离身份、风格与构图特征
  • 用掩码注意力机制精准控制生成过程
  • 无需提示词,支持跨域物体合成

基于参考图像的物体合成需将前景对象与背景场景融合生成协调图像。跨域场景下,模型需在保留对象身份的同时匹配风格化环境,这一难题目前被分割为缺乏生成保真度的“混合器”和需逐对象微调的“生成器”。本文提出 Insert In Style,首个零样本生成框架,兼具实用性与高保真。核心创新包括:(i) 多阶段训练协议解耦身份、风格与构图表征;(ii) 专用掩码注意力架构在生成中强制解耦;(iii) 先验保持目标维护已学身份与风格先验。该设计缓解统一注意力中的概念干扰,确保对多样化参考与风格的强泛化能力。模型在11.5万样本的新数据集上训练,数据来自结合大规模生成与迭代人机过滤的新型数据管道,保障语义身份保真与风格一致性。相较以往方法,本模型真正实现零样本,无需文本提示。我们还引入新公开基准用于风格化合成评估。实验表明,在身份与风格指标上均显著优于现有方法,用户研究结果有力支持该结论。

原文摘要 · Abstract (English)

Reference-based object composition involves integrating foreground reference image with background scene to produce harmonious fused image. This task becomes particularly challenging in cross-domain scenarios, where models must balance preserving the reference object's identity while harmonizing them to match stylized environments. This under-explored problem is currently split between practical "blenders" that lack generative fidelity and "generators" that require impractical, per-subject online finetuning. In this work, we introduce Insert In Style, the first zero-shot generative framework that is both practical and high-fidelity. Our core contribution is a unified framework with two key innovations: (i) a novel multi-stage training protocol that disentangles representations for identity, style, and composition, and (ii) a specialized masked-attention architecture that surgically enforces this disentanglement during generation (iii) A prior preservation objective that keeps learned identity and style priors intact. By design, this approach mitigates concept interference typical in unified-attention architectures while ensuring robust generalization across diverse references and styles. Our framework is trained on a new 115k sample dataset, curated from a novel data pipeline. This pipeline couples large-scale generation with a rigorous, iterative human-in-the-loop filtering process to ensure both high-fidelity semantic identity and style coherence. Unlike prior work, our model is truly zero-shot and requires no text prompts. We also introduce a new public benchmark for stylized composition. We demonstrate state-of-the-art performance, significantly outperforming existing methods on both identity and style metrics, a result strongly corroborated by user studies.

图像合成零样本风格迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。