arXiv:2605.19750cs.CV2026-05

解决视觉自回归模型持续个性化生成中的遗忘与概念纠缠问题。

CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models

论文配图:CPC-VAR:Continual Personalized and Compositional Generation in Visual Autoregressive Models
图 1 · 摘自论文原文
  • 通过选择性约束关键神经元,避免单概念学习时的灾难性遗忘。
  • 利用空间条件引导多分支特征融合,实现多概念图像的精准组合。
  • 适合需要长期更新个性化图像生成能力的研究者与应用开发。

视觉自回归(VAR)模型在文本到图像生成中展现出高效潜力。然而,现有基于VAR的个性化方法仍局限于静态场景,无法适应用户需求的动态演变。具体表现为:顺序概念学习导致严重灾难性遗忘,多概念合成常出现特征纠缠和属性不一致。本文首次系统研究了VAR模型中的持续个性化生成问题,识别出两大挑战:(i)在序列定制过程中保持已学概念;(ii)可控地组合多个个性化概念。为此,提出统一框架,包含两个核心组件:针对持续单概念学习,设计基于梯度的概念神经元选择(GCNS),仅约束任务间冲突参数,有效缓解遗忘且无需扩展模型;针对多概念合成,提出上下文感知的组合策略,通过多分支特征建模与局部交叉注意力融合,结合空间条件引导,实现精确且解耦的概念组合。大量实验表明,该方法在长序列持续个性化生成中显著提升性能,并在多概念图像合成上优于现有基线。结果证明了VAR模型在可扩展、可控个性化生成方面的潜力。

原文摘要 · Abstract (English)

Visual autoregressive (VAR) models have recently emerged as an efficient paradigm for text-to-image generation. Despite their strong generative capability, existing VAR-based personalization methods remain limited to static settings, failing to accommodate evolving user demands. In particular, sequential concept learning leads to severe catastrophic forgetting, while multi-concept synthesis often suffers from feature entanglement and attribute inconsistency. In this work, we present the first systematic study of continual personalized generation in VAR models. We identify two key challenges: (i) preserving previously learned concepts during sequential customization, and (ii) composing multiple personalized concepts in a controllable manner. To address these issues, we propose a unified framework with two core components. For continual single-concept learning, we introduce Gradient-based Concept Neuron Selection (GCNS), which identifies concept-relevant neurons and constrains only conflicting parameters across tasks, effectively mitigating forgetting without additional model expansion. For multi-concept synthesis, we propose a context-aware composition strategy that performs multi-branch feature modeling and localized cross-attention fusion guided by spatial conditions, enabling precise and disentangled concept composition. Extensive experiments demonstrate that our method significantly improves performance in long-sequence continual personalization while achieving superior results in multi-concept image synthesis compared to existing baselines. These findings highlight the potential of VAR models for scalable and controllable personalized generation.

个性化生成持续学习图像合成自回归模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。