解决少样本异常布局生成中的语义破碎问题,提升图像生成质量。
Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation

- 分离语义与视觉元素,用锚点稳定身份,用可重组构件建模细节。
- 5次示例下生成图像在视觉质量和布局对齐上均超越现有方法。
- 适合需要精细控制图像生成的少样本场景,如设计、艺术创作。
布局到图像(L2I)任务通过对象类别和空间布局实现图像生成的细粒度控制。然而,现有L2I方法在少样本异常设置下会产生碎片化和失真的生成结果。我们称此现象为表示碎片化,源于语义身份与视觉细节之间的粒度不匹配。为此,我们提出一种以表示为核心的框架,将语义与基元解耦,实现鲁棒的少样本适应。具体而言,语义锚定将类别语义聚合为稳定的身份锚点,而基元赋予则建模可重构的基元以增强局部细节建模能力。概念引导进一步引入显著性感知目标,调节优化过程以保持前景语义一致性。大量实验表明,在5次示例条件下,本方法在多样化的异常领域中,均一致优于当前最优的L2I方法,在视觉保真度和布局对齐方面均有显著提升。源代码已公开于 https://github.com/iCVTEAM/DSP。
原文摘要 · Abstract (English)
The layout-to-image (L2I) task enables fine-grained control over image generation via object categories and spatial layouts. However, existing L2I methods yield fragmented and distorted generations under few-shot atypical settings. We term this failure as representation fragmentation, arising from a granularity mismatch that entangles semantic identity with visual details. To address this issue, we propose a representation-driven framework that disentangles semantics from primitives for robust few-shot adaptation. Specifically, Semantic Anchoring aggregates categorical semantics into anchors for stable identity, while Primitive Imbuing models recomposable primitives for robust local detail modeling. Conceptual Steering further regulates optimization with a saliency-aware objective to preserve foreground semantic consistency. Extensive experiments demonstrate consistent improvements in the 5-shot regime over state-of-the-art L2I methods in both visual fidelity and alignment across diverse atypical domains. The source code is publicly available at https://github.com/iCVTEAM/DSP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。