arXiv:2601.06965cs.CV2026-01被引 7

统一模型首次实现个人化理解、生成与编辑的端到端协同。

Unified Personalized Understanding, Generating and Editing

  • 用分离任务的结构化概念令牌减少跨任务干扰。
  • 通过显式知识回放实现个性化属性在多任务间一致传递。
  • 适合需要可控个人化交互的研究者与开发者。

统一大型多模态模型(LMMs)在通用多模态理解与生成方面取得显著进展,但仍受制于‘一刀切’范式,难以一致且可控地建模用户特定概念(如生成<maeve>的照片)。现有个性化方法多依赖外部检索,效率低且难以融入统一多模态流程。近期方法引入可学习软提示编码概念信息,但或耦合理解与生成,或依赖复杂多阶段训练,导致跨任务干扰,使个性化知识模糊或错位。本文提出OmniPersona,首个将个性化理解、生成与图像编辑统一于单一架构的端到端框架。该框架采用结构解耦的概念令牌,为不同任务分配专用子空间以最小化干扰,并引入显式知识回放机制,跨任务传播个性化属性知识,实现一致行为。为系统评估统一个性化,我们构建OmniPBench,扩展UnifyBench概念集,加入个性化编辑任务与跨任务评估协议。实验表明,OmniPersona在多样个性化任务中表现优异且稳健。我们期望其成为可控统一个性化研究的有力基线。

原文摘要 · Abstract (English)

Unified large multimodal models (LMMs) have achieved remarkable progress in general-purpose multimodal understanding and generation. However, they still operate under a ``one-size-fits-all'' paradigm and struggle to model user-specific concepts (e.g., generate a photo of \texttt{<maeve>}) in a consistent and controllable manner. Existing personalization methods typically rely on external retrieval, which is inefficient and poorly integrated into unified multimodal pipelines. Recent personalized unified models introduce learnable soft prompts to encode concept information, yet they either couple understanding and generation or depend on complex multi-stage training, leading to cross-task interference and ultimately to fuzzy or misaligned personalized knowledge. We present \textbf{OmniPersona}, an end-to-end personalization framework for unified LMMs that, for the first time, integrates personalized understanding, generation, and image editing within a single architecture. OmniPersona introduces structurally decoupled concept tokens, allocating dedicated subspaces for different tasks to minimize interference, and incorporates an explicit knowledge replay mechanism that propagates personalized attribute knowledge across tasks, enabling consistent personalized behavior. To systematically evaluate unified personalization, we propose \textbf{\texttt{OmniPBench}}, extending the public UnifyBench concept set with personalized editing tasks and cross-task evaluation protocols integrating understanding, generation, and editing. Experimental results demonstrate that OmniPersona delivers competitive and robust performance across diverse personalization tasks. We hope OmniPersona will serve as a strong baseline and spur further research on controllable, unified personalization.

多模态个性化统一模型知识回放

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。