arXiv:2603.12743cs.CVcs.AI2026-03被引 1

让模型把文字知识精准映射到视觉概念,提升罕见词生成质量。

MoKus: Leveraging Cross-Modal Knowledge Transfer for Knowledge-Aware Concept Customization

  • 通过跨模态知识迁移,将文本知识自动同步到视觉生成中。
  • 在新基准KnowCusBench上显著优于现有方法,生成更忠实于知识的图像。
  • 适用于虚拟概念创建、概念擦除等下游任务,扩展性强。

概念定制通常将稀有词绑定到目标视觉概念,但预训练数据中缺乏这些稀有词,导致性能不稳定,且无法传递概念本质知识。为此,我们提出知识感知的概念定制任务,旨在将多样化文本知识绑定至目标视觉概念,要求模型识别提示中的知识并实现高保真定制化生成。为此,我们提出MoKus框架:第一阶段学习目标概念的锚定表示以存储视觉信息;第二阶段通过更新知识查询的答案至锚定表示,实现知识与概念的高效绑定。为全面评估,我们构建首个知识感知概念定制基准KnowCusBench。大量实验表明,MoKus优于当前最优方法。跨模态知识迁移机制使其可轻松扩展至虚拟概念生成、概念擦除等应用,并在世界知识基准上实现性能提升。

原文摘要 · Abstract (English)

Concept customization typically binds rare tokens to a target concept. Unfortunately, these approaches often suffer from unstable performance as the pretraining data seldom contains these rare tokens. Meanwhile, these rare tokens fail to convey the inherent knowledge of the target concept. Consequently, we introduce Knowledge-aware Concept Customization, a novel task aiming at binding diverse textual knowledge to target visual concepts. This task requires the model to identify the knowledge within the text prompt to perform high-fidelity customized generation. Meanwhile, the model should efficiently bind all the textual knowledge to the target concept. Therefore, we propose MoKus, a novel framework for knowledge-aware concept customization. Our framework relies on a key observation: cross-modal knowledge transfer, where modifying knowledge within the text modality naturally transfers to the visual modality during generation. Inspired by this observation, MoKus contains two stages: (1) In visual concept learning, we first learn the anchor representation to store the visual information of the target concept. (2) In textual knowledge updating, we update the answer for the knowledge queries to the anchor representation, enabling high-fidelity customized generation. To further comprehensively evaluate our proposed MoKus on the new task, we introduce the first benchmark for knowledge-aware concept customization: KnowCusBench. Extensive evaluations have demonstrated that MoKus outperforms state-of-the-art methods. Moreover, the cross-model knowledge transfer allows MoKus to be easily extended to other knowledge-aware applications like virtual concept creation and concept erasure. We also demonstrate the capability of our method to achieve improvements on world knowledge benchmarks.

概念定制跨模态知识迁移图像生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。