用贝叶斯优化交互式合并图像生成模型,提升风格混合效率。
GimmBO: Interactive Generative Image Model Merging via Bayesian Optimization
- 通过偏好贝叶斯优化自动搜索最佳模型权重组合
- 在20-30个适配器中实现更快收敛与更高成功率
- 适合需要快速探索视觉风格混合的设计师和研究者
基于微调的适应方法广泛用于定制扩散模型图像生成,催生了大量社区创建的适配器,涵盖多样主题与风格。相同基模型的适配器可通过加权合并,生成广阔连续的设计空间。现有工作依赖手动滑块调节,难以扩展,即便候选适配器仅20-30个也难高效选择。本文提出GimmBO,基于偏好贝叶斯优化(PBO)支持交互式适配器合并探索。针对真实使用中的稀疏性与权重范围受限现象,设计两阶段贝叶斯优化后端,提升高维空间下的采样效率与收敛速度。通过模拟用户与真实用户研究验证,本方法在收敛速度、成功率达85%以上,显著优于标准贝叶斯优化与线性搜索基线,并展示了框架的可扩展性。
原文摘要 · Abstract (English)
Fine-tuning-based adaptation is widely used to customize diffusion-based image generation, leading to large collections of community-created adapters that capture diverse subjects and styles. Adapters derived from the same base model can be merged with weights, enabling the synthesis of new visual results within a vast and continuous design space. To explore this space, current workflows rely on manual slider-based tuning, an approach that scales poorly and makes weight selection difficult, even when the candidate set is limited to 20-30 adapters. We propose GimmBO to support interactive exploration of adapter merging for image generation through Preferential Bayesian Optimization (PBO). Motivated by observations from real-world usage, including sparsity and constrained weight ranges, we introduce a two-stage BO backend that improves sampling efficiency and convergence in high-dimensional spaces. We evaluate our approach with simulated users and a user study, demonstrating improved convergence, high success rates, and consistent gains over BO and line-search baselines, and further show the flexibility of the framework through several extensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。