arXiv:2605.22743cs.LG2026-05

SeqLoRA提升多概念图像生成的稳定性与可扩展性

SeqLoRA: Bilevel Orthogonal Adaptation for Continual Multi-Concept Generation

论文配图:SeqLoRA: Bilevel Orthogonal Adaptation for Continual Multi-Concept Generation
图 1 · 摘自论文原文
  • 通过双层优化联合训练LoRA两个因子,避免概念干扰
  • 支持最多101个概念,生成图像身份保持更准确
  • 无需昂贵融合步骤,适合个性化图像生成场景

参数高效微调可快速定制文生图扩散模型,但多概念组合仍面临表征干扰难题。现有模块化方法要么依赖昂贵的后期融合,要么冻结适配子空间,限制表达能力与概念保真度。为此,我们提出序列正则化LoRA(SeqLoRA),一种约束连续学习框架,通过双层优化联合优化LoRA两个因子。理论上,我们建立了算法的强收敛性保证,并将残差层激活建模为矩阵亚高斯过程,推导出灾难性遗忘的高概率上界。进一步证明,从数据中学习LoRA基比固定基方法更有效降低残差干扰能量。在多概念图像生成实验中,SeqLoRA在最多101个概念下提升了身份保持能力和可扩展性,同时避免了高昂融合操作,减少了组合生成中的属性干扰。

原文摘要 · Abstract (English)

Parameter-efficient fine-tuning enables fast personalization of text-to-image diffusion models, but composing multiple custom concepts remains challenging due to representation interference. Existing modular methods either rely on expensive post-hoc fusion or freeze adaptation subspaces, which limit expressiveness and concept fidelity. To address this trade-off, we propose Sequential regularized LoRA (SeqLoRA), a constrained continual learning framework that jointly optimizes both LoRA factors via bilevel optimization. Theoretically, we establish strong convergence guarantees for our algorithm and model the residual layer activations as a matrix sub-Gaussian process to derive high-probability bounds on catastrophic forgetting. We further prove that learning the LoRA basis from data minimizes residual interference energy more effectively than frozen-basis methods. Experiments on multi-concept image generation demonstrate that SeqLoRA improves identity preservation and scalability across up to 101 concepts, while avoiding costly fusion and reducing attribute interference in composed generations.

图像生成LoRA持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。