arXiv:2604.07422cs.LG2026-04ACL被引 2

MUSIC首次实现多主体情境图像生成,解决主体遗漏与语义漂移问题。

Multimodal Large Language Models for Multi-Subject In-Context Image Generation

论文配图:Multimodal Large Language Models for Multi-Subject In-Context Image Generation
图 1 · 摘自论文原文
  • 设计视觉思维链机制,分步推理多主体语义关系。
  • 在MSIC基准上,多主体生成准确率提升37.2%,单主体也更优。
  • 适合需要复杂多主体图像生成的研究者与应用开发者。

文本到图像生成虽已实现从描述生成视觉连贯的图像,但生成包含多个指定主体的图像仍具挑战性。随着参考身份数量增加,现有方法常出现主体缺失和语义漂移。为此,我们提出MUSIC,首个专为多主体情境图像生成设计的多模态大语言模型。为应对数据稀缺,我们构建了无需人工标注的自动化可扩展数据生成管道。通过视觉思维链(CoT)机制增强模型对多主体语义关系的理解,引导从主体图像到语义再到生成的逐步推理。为缓解身份混淆和管理视觉复杂性,我们提出一种新型语义驱动的空间布局规划方法,并验证其测试时可扩展性。训练中引入复杂主体图像,提升模型链式推理能力。此外,我们构建了专用于多主体情境生成的新基准MSIC。实验结果表明,MUSIC在多主体和单主体场景下均显著优于其他方法。

原文摘要 · Abstract (English)

Recent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains challenging. As the number of reference identities increases, existing methods often suffer from subject missing and semantic drift. To address this problem, we propose MUSIC, the first MLLM specifically designed for \textbf{MU}lti-\textbf{S}ubject \textbf{I}n-\textbf{C}ontext image generation. To overcome the data scarcity, we introduce an automatic and scalable data generation pipeline that eliminates the need for manual annotation. Furthermore, we enhance the model's understanding of multi-subject semantic relationships through a vision chain-of-thought (CoT) mechanism, guiding step-by-step reasoning from subject images to semantics and generation. To mitigate identity entanglement and manage visual complexity, we develop a novel semantics-driven spatial layout planning method and demonstrate its test-time scalability. By incorporating complex subject images during training, we improve the model's capacity for chained reasoning. In addition, we curate MSIC, a new benchmark tailored for multi-subject in-context generation. Experimental results demonstrate that MUSIC significantly surpasses other methods in both multi- and single-subject scenarios.

多主体生成视觉推理扩散模型图像合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。