arXiv:2603.20470cs.AI2026-03被引 4

自动整合网络专家模型,按需组合生成图像。

DiffGraph: An Automated Agent-driven Model Merging Framework for In-the-Wild Text-to-Image Generation

  • 构建图结构自动管理在线专家模型,支持动态扩展。
  • 根据用户需求激活子图,灵活组合专家实现定制生成。
  • 适用于需要多样化图像生成的开放场景,适合研究者和开发者。

文本到图像(T2I)社区的快速发展催生了大量在线专家模型,这些模型是预训练扩散模型的变体,具备多样化的生成能力。然而,现有模型融合方法难以充分利用丰富的在线专家资源,且无法满足多样的真实场景用户需求。我们提出DiffGraph,一种基于代理的图结构模型融合框架,可自动获取在线专家并灵活融合以应对不同用户需求。DiffGraph通过节点注册与校准构建可扩展的图结构,将不断增长的在线专家组织其中。随后,根据用户需求动态激活特定子图,实现不同专家的灵活组合,达成用户期望的生成效果。大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

The rapid growth of the text-to-image (T2I) community has fostered a thriving online ecosystem of expert models, which are variants of pretrained diffusion models specialized for diverse generative abilities. Yet, existing model merging methods remain limited in fully leveraging abundant online expert resources and still struggle to meet diverse in-the-wild user needs. We present DiffGraph, a novel agent-driven graph-based model merging framework, which automatically harnesses online experts and flexibly merges them for diverse user needs. Our DiffGraph constructs a scalable graph and organizes ever-expanding online experts within it through node registration and calibration. Then, DiffGraph dynamically activates specific subgraphs based on user needs, enabling flexible combinations of different experts to achieve user-desired generation. Extensive experiments show the efficacy of our method.

模型融合图像生成自动化扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。