arXiv:2605.31266cs.CVcs.AI2026-05中稿 · ICML被引 1

解决少样本异常布局生成中的语义破碎问题,提升图像生成质量。

Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation

论文配图:Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation
图 1 · 摘自论文原文
  • 分离语义与视觉元素,用锚点稳定身份,用可重组构件建模细节。
  • 5次示例下生成图像在视觉质量和布局对齐上均超越现有方法。
  • 适合需要精细控制图像生成的少样本场景,如设计、艺术创作。

布局到图像(L2I)任务通过对象类别和空间布局实现图像生成的细粒度控制。然而,现有L2I方法在少样本异常设置下会产生碎片化和失真的生成结果。我们称此现象为表示碎片化,源于语义身份与视觉细节之间的粒度不匹配。为此,我们提出一种以表示为核心的框架,将语义与基元解耦,实现鲁棒的少样本适应。具体而言,语义锚定将类别语义聚合为稳定的身份锚点,而基元赋予则建模可重构的基元以增强局部细节建模能力。概念引导进一步引入显著性感知目标,调节优化过程以保持前景语义一致性。大量实验表明,在5次示例条件下,本方法在多样化的异常领域中,均一致优于当前最优的L2I方法,在视觉保真度和布局对齐方面均有显著提升。源代码已公开于 https://github.com/iCVTEAM/DSP。

原文摘要 · Abstract (English)

The layout-to-image (L2I) task enables fine-grained control over image generation via object categories and spatial layouts. However, existing L2I methods yield fragmented and distorted generations under few-shot atypical settings. We term this failure as representation fragmentation, arising from a granularity mismatch that entangles semantic identity with visual details. To address this issue, we propose a representation-driven framework that disentangles semantics from primitives for robust few-shot adaptation. Specifically, Semantic Anchoring aggregates categorical semantics into anchors for stable identity, while Primitive Imbuing models recomposable primitives for robust local detail modeling. Conceptual Steering further regulates optimization with a saliency-aware objective to preserve foreground semantic consistency. Extensive experiments demonstrate consistent improvements in the 5-shot regime over state-of-the-art L2I methods in both visual fidelity and alignment across diverse atypical domains. The source code is publicly available at https://github.com/iCVTEAM/DSP.

布局生成少样本学习语义解耦图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。