arXiv:2510.18083cs.CV2025-10被引 2

让模型按文字指令组合多图零件,生成新物体。

Chimera: Compositional Image Generation using Part-based Concepting

  • 用部件级语义原子构建数据集,实现零件级控制生成
  • 在37000个提示下训练,零件对齐准确率提升14%
  • 适合需要精细物体合成的创意设计与个性化生成场景

个性化图像生成模型虽能根据文本或单张图像生成图像,但难以在无用户标注的情况下,从多张源图中精确组合特定部件。为此,我们提出Chimera,一种基于部件概念的组合式图像生成模型,可根据文本指令将不同源图中的指定部件组合成新物体。我们首先基于464个独特的(部件,主体)组合构建分类体系,称为语义原子,并生成37,000个提示,利用高保真文生图模型合成对应图像。训练过程中,采用定制的扩散先验模型,结合部件条件引导,以强化语义身份与空间布局的一致性。我们还引入指标PartEval,用于评估生成管道的保真度与组合准确性。人评与该指标结果表明,Chimera在部件对齐与组合准确性上比基线高出14%,视觉质量提升21%。

原文摘要 · Abstract (English)

Personalized image generative models are highly proficient at synthesizing images from text or a single image, yet they lack explicit control for composing objects from specific parts of multiple source images without user specified masks or annotations. To address this, we introduce Chimera, a personalized image generation model that generates novel objects by combining specified parts from different source images according to textual instructions. To train our model, we first construct a dataset from a taxonomy built on 464 unique (part, subject) pairs, which we term semantic atoms. From this, we generate 37k prompts and synthesize the corresponding images with a high-fidelity text-to-image model. We train a custom diffusion prior model with part-conditional guidance, which steers the image-conditioning features to enforce both semantic identity and spatial layout. We also introduce an objective metric PartEval to assess the fidelity and compositional accuracy of generation pipelines. Human evaluations and our proposed metric show that Chimera outperforms other baselines by 14% in part alignment and compositional accuracy and 21% in visual quality.

图像生成部件组合扩散模型个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。