用超网络快速合并LoRA,实现手机端高效个性化图像生成
LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image Generation
- 通过预训练超网络学习跨内容风格的LoRA合并策略
- 合并速度提升4000倍以上,图像质量显著优于现有方法
- 结合多模态大模型评估,更适合实际应用与新场景
近期图像生成模型已支持用户自定义主体(内容)和风格的个性化图像生成。以往方法通过优化方式合并低秩适配器(LoRAs),计算开销大,不适用于资源受限设备如智能手机。为此,我们提出LoRA.rar,不仅提升图像质量,还将合并速度提升超过4000倍。我们构建了包含风格与主体LoRA的数据集,并在多样化的内容-风格LoRA对上预训练超网络,学习出可泛化至未见内容-风格组合的高效合并策略,实现实时、高质量个性化生成。此外,我们指出现有评价指标在内容-风格一致性上的不足,提出基于多模态大语言模型(MLLMs)的新评估协议。实验表明,该方法在内容与风格保真度上显著优于当前最先进水平,经由MLLM与人工评估验证。
原文摘要 · Abstract (English)
Recent advancements in image generation models have enabled personalized image creation with both user-defined subjects (content) and styles. Prior works achieved personalization by merging corresponding low-rank adapters (LoRAs) through optimization-based methods, which are computationally demanding and unsuitable for real-time use on resource-constrained devices like smartphones. To address this, we introduce LoRA$.$rar, a method that not only improves image quality but also achieves a remarkable speedup of over $4000\times$ in the merging process. We collect a dataset of style and subject LoRAs and pre-train a hypernetwork on a diverse set of content-style LoRA pairs, learning an efficient merging strategy that generalizes to new, unseen content-style pairs, enabling fast, high-quality personalization. Moreover, we identify limitations in existing evaluation metrics for content-style quality and propose a new protocol using multimodal large language models (MLLMs) for more accurate assessment. Our method significantly outperforms the current state of the art in both content and style fidelity, as validated by MLLM assessments and human evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。