arXiv:2509.18699cs.CV2025-09SIGGRAPH被引 5

提出AGSwap方法,实现跨类别物体融合的精准生成。

AGSwap: Overcoming Category Boundaries in Object Fusion via Adaptive Group Swapping

  • 通过分组嵌入交换与自适应更新机制融合不同概念特征
  • 在45万组融合对上超越现有最优方法,生成更一致图像
  • 适合虚拟现实、游戏等需复杂物体合成的场景

将跨类别物体融合为单一连贯对象在文本到图像生成中日益受到关注,广泛应用于虚拟现实、数字媒体、影视和游戏等领域。然而,现有方法常因重叠伪影和整合不佳导致结果偏倚、视觉混乱或语义不一致。此外,该领域进展受限于缺乏全面的基准数据集。为此,我们提出 extbf{自适应分组交换(AGSwap)},一种简单但高效的方案,包含两个关键组件:(1) 分组嵌入交换,通过特征操作融合不同概念的语义属性;(2) 自适应分组更新,基于平衡评估分数动态优化,确保合成一致性。同时,我们构建了 extbf{跨类别物体融合(COF)}数据集,基于ImageNet-1K和WordNet,包含95个超类,每类10个子类,共支持451,250种唯一融合对。大量实验表明,AGSwap在使用简单与复杂提示时均优于当前最先进的组合式文本到图像方法,包括GPT-Image-1。

原文摘要 · Abstract (English)

Fusing cross-category objects to a single coherent object has gained increasing attention in text-to-image (T2I) generation due to its broad applications in virtual reality, digital media, film, and gaming. However, existing methods often produce biased, visually chaotic, or semantically inconsistent results due to overlapping artifacts and poor integration. Moreover, progress in this field has been limited by the absence of a comprehensive benchmark dataset. To address these problems, we propose \textbf{Adaptive Group Swapping (AGSwap)}, a simple yet highly effective approach comprising two key components: (1) Group-wise Embedding Swapping, which fuses semantic attributes from different concepts through feature manipulation, and (2) Adaptive Group Updating, a dynamic optimization mechanism guided by a balance evaluation score to ensure coherent synthesis. Additionally, we introduce \textbf{Cross-category Object Fusion (COF)}, a large-scale, hierarchically structured dataset built upon ImageNet-1K and WordNet. COF includes 95 superclasses, each with 10 subclasses, enabling 451,250 unique fusion pairs. Extensive experiments demonstrate that AGSwap outperforms state-of-the-art compositional T2I methods, including GPT-Image-1 using simple and complex prompts.

图像生成物体融合文本生成图像AI艺术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。