arXiv:2512.12675cs.CVcs.AI2025-12被引 9

让图像生成同时处理多个主体并准确区分它们

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling

  • 用统一模型同时实现主体组合与区分
  • 在两个基准上超越现有开源模型表现
  • 适合需要精准控制多主体图像的场景

主体驱动的图像生成已从单主体发展到多主体组合,但忽略了区分能力——即当输入包含多个候选主体时,正确识别并生成对应主体的能力。这一缺陷限制了其在复杂真实视觉场景中的应用。本文提出Scone,一种融合组合与区分的统一理解-生成方法。通过让理解专家作为语义桥梁,将语义信息传递给生成专家,以保持主体身份并减少干扰。采用两阶段训练:先学习组合,再通过语义对齐和基于注意力的掩码增强区分能力。我们还构建了SconeEval,一个用于评估多种场景下组合与区分能力的基准。实验表明,Scone在两个基准上均优于现有开源模型。模型、基准与训练数据已公开于:https://github.com/Ryann-Ran/Scone。

原文摘要 · Abstract (English)

Subject-driven image generation has advanced from single- to multi-subject composition, while neglecting distinction, the ability to distinguish and generate the correct subject when inputs contain multiple candidates. This limitation restricts effectiveness in complex, realistic visual settings. We propose Scone, a unified understanding-generation method that integrates composition and distinction. Scone enables the understanding expert to act as a semantic bridge, conveying semantic information and guiding the generation expert to preserve subject identity while minimizing interference. A two-stage training scheme first learns composition, then enhances distinction through semantic alignment and attention-based masking. We also introduce SconeEval, a benchmark for evaluating both composition and distinction across diverse scenarios. Experiments demonstrate that Scone outperforms existing open-source models in composition and distinction tasks on two benchmarks. Our model, benchmark, and training data are available at: https://github.com/Ryann-Ran/Scone.

图像生成多主体语义区分

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。