arXiv:2603.21136cs.CV2026-03

让用户自由定义多个主体的排列和位置,生成精准可控的多主体图像。

MS-CustomNet: Controllable Multi-Subject Customization with Hierarchical Relational Semantics

  • 通过分层关系语义建模,实现多主体布局的显式控制。
  • 身份保留度达0.61,位置控制精度达0.94,优于现有方法。
  • 适用于需要精细调控多对象场景的创作与设计任务。

基于扩散的文本到图像生成已取得显著进展,但在保持主体间精细交互的前提下对多主体场景进行定制仍具挑战。现有方法难以提供对组合结构和主体间空间关系的明确用户控制。为此,我们提出MS-CustomNet,一种新型多主体定制框架。该框架支持零样本集成多个用户提供的对象,并能显式定义其层级排列与空间位置。方法在保持个体主体身份的同时,学习并实现用户指定的主体间组合关系。我们还构建了源自COCO的MSI数据集,以支持此类复杂多主体组合的训练。实验表明,MS-CustomNet在多主体定制任务中实现0.61的DINO-I身份保留得分与0.94的YOLO-L位置控制得分,展现出生成高保真图像并实现用户导向的精确多主体构图与空间控制的卓越能力。

原文摘要 · Abstract (English)

Diffusion-based text-to-image generation has advanced significantly, yet customizing scenes with multiple distinct subjects while maintaining fine-grained control over their interactions remains challenging. Existing methods often struggle to provide explicit user-defined control over the compositional structure and precise spatial relationships between subjects. To address this, we introduce MS-CustomNet, a novel framework for multi-subject customization. MS-CustomNet allows zero-shot integration of multiple user-provided objects and, crucially, empowers users to explicitly define these hierarchical arrangements and spatial placements within the generated image. Our approach ensures individual subject identity preservation while learning and enacting these user-specified inter-subject compositions. We also present the MSI dataset, derived from COCO, to facilitate training on such complex multi-subject compositions. MS-CustomNet offers enhanced, fine-grained control over multi-subject image generation. Our method achieves a DINO-I score of 0.61 for identity preservation and a YOLO-L score of 0.94 for positional control in multi-subject customization tasks, demonstrating its superior capability in generating high-fidelity images with precise, user-directed multi-subject compositions and spatial control.

多主体生成空间控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。