arXiv:2509.04126cs.CVcs.AI2025-09

让复杂图文生成更精准多样,通过专家分工协作提升细节与风格表现

MEPG:Multi-Expert Planning and Generation for Compositionally-Rich Image Generation

  • 用语言模型拆解提示词为位置和风格指令,实现精准控制
  • 动态调度不同专家模型,局部区域生成质量与风格显著提升
  • 支持实时布局修改和风格切换,适合需要精细调控的创作场景

文本到图像扩散模型虽已达到优异图像质量,但在处理多元素复杂提示词及风格多样性方面仍存在局限。为此,我们提出多专家规划与生成框架(MEPG),通过位置与风格感知的大语言模型(LLM)与空间语义专家模块协同工作。该框架包含两个核心组件:(1) 位置-风格感知(PSA)模块,利用监督微调的LLM将输入提示分解为精确的空间坐标与风格化语义指令;(2) 多专家扩散(MED)模块,通过跨区域动态专家路由实现全局与局部区域的生成。在每个局部区域生成时,基于注意力门控机制选择性激活特定专家模型(如写实专家、风格化专家)。架构支持轻量级专家替换与扩展,具备强可拓展性。同时,交互式界面支持实时空间布局编辑与分区域风格选择。实验表明,MEPG在相同骨干模型下,显著优于基线,在图像质量与风格多样性上均有明显提升。

原文摘要 · Abstract (English)

Text-to-image diffusion models have achieved remarkable image quality, but they still struggle with complex, multiele ment prompts, and limited stylistic diversity. To address these limitations, we propose a Multi-Expert Planning and Gen eration Framework (MEPG) that synergistically integrates position- and style-aware large language models (LLMs) with spatial-semantic expert modules. The framework comprises two core components: (1) a Position-Style-Aware (PSA) module that utilizes a supervised fine-tuned LLM to decom pose input prompts into precise spatial coordinates and style encoded semantic instructions; and (2) a Multi-Expert Dif fusion (MED) module that implements cross-region genera tion through dynamic expert routing across both local regions and global areas. During the generation process for each lo cal region, specialized models (e.g., realism experts, styliza tion specialists) are selectively activated for each spatial par tition via attention-based gating mechanisms. The architec ture supports lightweight integration and replacement of ex pert models, providing strong extensibility. Additionally, an interactive interface enables real-time spatial layout editing and per-region style selection from a portfolio of experts. Ex periments show that MEPG significantly outperforms base line models with the same backbone in both image quality and style diversity.

图像生成多专家扩散模型风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。