让文字描述精准生成多个物体,还能灵活控制位置和属性。
MoGen: A Unified Collaborative Framework for Controllable Multi-Object Image Generation
- 用语义锚点模块把文字描述对准图像区域
- 生成物体数量准确,质量优于现有方法
- 适合需要精细控制的创作人群
现有多物体图像生成方法难以实现语言描述与图像区域的精确对齐,常导致物体数量不一致或属性混淆。主流方法依赖外部控制信号约束空间布局和视觉属性,但输入格式僵化,难以适配用户多样需求。为此,我们提出MoGen,一种用户友好的多物体图像生成框架。首先设计区域语义锚点(RSA)模块,在生成过程中将语言描述中的短语单位精确锚定到对应图像区域,实现按数量要求生成多物体。在此基础上,引入自适应多模态引导(AMG)模块,动态解析并融合多种来源的控制信号,形成结构化意图,进而选择性地指导场景布局与物体属性约束,实现动态细粒度控制。实验表明,MoGen在生成质量、数量一致性与细粒度控制方面显著优于现有方法,同时具备更高可用性与控制灵活性。代码已开源:https://github.com/Tear-kitty/MoGen/tree/master。
原文摘要 · Abstract (English)
Existing multi-object image generation methods face difficulties in achieving precise alignment between localized image generation regions and their corresponding semantics based on language descriptions, frequently resulting in inconsistent object quantities and attribute aliasing. To mitigate this limitation, mainstream approaches typically rely on external control signals to explicitly constrain the spatial layout, local semantic and visual attributes of images. However, this strong dependency makes the input format rigid, rendering it incompatible with the heterogeneous resource conditions of users and diverse constraint requirements. To address these challenges, we propose MoGen, a user-friendly multi-object image generation method. First, we design a Regional Semantic Anchor (RSA) module that precisely anchors phrase units in language descriptions to their corresponding image regions during the generation process, enabling text-to-image generation that follows quantity specifications for multiple objects. Building upon this foundation, we further introduce an Adaptive Multi-modal Guidance (AMG) module, which adaptively parses and integrates various combinations of multi-source control signals to formulate corresponding structured intent. This intent subsequently guides selective constraints on scene layouts and object attributes, achieving dynamic fine-grained control. Experimental results demonstrate that MoGen significantly outperforms existing methods in generation quality, quantity consistency, and fine-grained control, while exhibiting superior accessibility and control flexibility. Code is available at: https://github.com/Tear-kitty/MoGen/tree/master.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。