arXiv:2605.20237cs.CV2026-05

用单图控制动漫角色生成,保持外观一致且无需微调。

AnimeAdapter: A Modular Adapter for Appearance-Consistent Anime Character Generation

论文配图:AnimeAdapter: A Modular Adapter for Appearance-Consistent Anime Character Generation
图 1 · 摘自论文原文
  • 通过局部注意力机制注入参考图的细节特征。
  • 支持姿态变化下的角色外观一致性,生成效果稳定。
  • 轻量模块化设计,兼容现有扩散模型流程。

我们提出一种轻量级外观适配器,用于Stable Diffusion,可在多种编辑条件下实现可控且一致的动漫角色生成。该方法不依赖大规模视觉语言模型或针对个体的微调,而是将单张参考图像中的细粒度视觉特征注入扩散过程。基于CLIP的隐式局部空间化,我们开发了语义选择性局部注意力机制;为进一步解耦角色外观与空间布局,训练时引入姿态感知条件。所获预训练适配器体积小、模块化,完全兼容Stable Diffusion社区工作流,部署时无需额外微调。此外,我们构建了一个基于精心筛选与重构Danbooru提示词的高质量动漫角色数据集,并在多个实际角色编辑场景中评估方法性能。代码、模型权重及数据集将在论文录用后公开发布。

原文摘要 · Abstract (English)

We present a lightweight appearance adapter for Stable Diffusion that enables controllable and consistent anime character generation under diverse editing conditions. Instead of relying on large-scale vision-language models or per-subject fine-tuning, our method injects fine-grained visual features from a single reference image into the diffusion process. Based on CLIP emergent local spatialization, we develop semantic-selective local attention. To further disentangle character appearance from spatial layout, we incorporate pose-aware conditioning during adapter training. The resulting pretrained adapter remains compact, modular, and fully compatible with Stable Diffusion community workflows, while requiring no additional fine-tuning at deployment time. Furthermore, we present a high-quality anime character dataset based on curated and restructured Danbooru prompts, and evaluate our method across several practical character editing scenarios. Our code, model weights, and dataset will be publicly released upon acceptance.

动漫生成扩散模型图像控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。