arXiv:2508.07341cs.CV2025-08被引 1

让预训练模型深度注入新概念,实现个性化图像生成

DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image Generation

  • 通过分层多模态上下文学习,深度融合新概念
  • 仅用少量参数就达到适配方法的生成效果
  • 支持零训练风格定制,适合快速个性化应用

统一自回归(AR)模型在多模态理解与生成方面表现优异,但在个性化图像生成领域尚未充分发挥潜力。现有定制化方法面临根本矛盾:基于适配的方法易过拟合且扩展性差,而概念注入范式因浅层注入策略导致视觉保真度低、重上下文能力弱。为此,我们提出DCoAR,一种全新的深度概念注入框架,完全冻结预训练模型。DCoAR通过分层多模态上下文学习(LMCL)策略深度整合新概念,并采用多维度正则化机制稳定训练:双重先验保持(DPP)损失缓解语义漂移,上下文感知自正则化(CASR)损失增强重上下文能力。该框架还支持用户提供的风格下无训练主体定制。实验表明,DCoAR显著优于以往注入方法,性能接近适配方法,但所需可训练参数大幅减少。

原文摘要 · Abstract (English)

The unified autoregressive (AR) model excels at multimodal understanding and generation. However, its full potential in the domain of customized image generation has yet to be fully realized. Existing customization approaches for unified AR models face a fundamental dilemma: adaptation-based methods suffer from overfitting and scalability bottlenecks, while concept-injection paradigms are constrained by a shallow injection strategy that leads to poor visual fidelity and impaired re-contextualization. To address this, we propose DCoAR, a novel deep concept injection framework that maintains a completely frozen pre-trained model. DCoAR deeply integrates new concepts through a Layer-wise Multimodal Context Learning (LMCL) strategy, which is stabilized by a multi-faceted regularization scheme: a Dual Prior Preservation (DPP) loss to mitigate semantic drift and a Context-Aware Self-Regularization (CASR) loss to enhance re-contextualization. The framework also enables training-free subject customization in user-provided styles. Experiments demonstrate that DCoAR significantly outperforms previous injection-based methods and achieves performance competitive with adaptation-based approaches while requiring substantially fewer trainable parameters. Code: https://github.com/KZF-kzf/CoAR

个性化生成概念注入自回归模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。