arXiv:2412.07589cs.CV2024-12CVPR被引 23

让漫画生成可定制角色,实现多角色动态控制。

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation

  • 用多模态大模型做角色特征适配器,结合扩散模型生成
  • 在43,264页漫画数据上训练,支持角色表情动作灵活调整
  • 适合需要精细角色控制的漫画创作与动画前期设计

故事可视化任务旨在从文本描述生成视觉叙事,现有文本到图像生成模型在多角色场景中对角色外观和互动缺乏有效控制。为此,我们提出新任务:定制化漫画生成,并引入DiffSensei框架,该框架将基于扩散的图像生成器与多模态大语言模型(MLLM)结合,后者作为文本兼容的身份适配器。通过掩码交叉注意力机制,实现角色特征的无缝融合,无需直接像素传递即可精准控制布局。MLLM适配器根据分镜文本线索动态调整角色特征,支持表情、姿态和动作的灵活变化。我们还构建了大规模数据集MangaZero,包含43,264页漫画和427,147个标注分镜,支持跨帧角色交互与运动的可视化。大量实验表明,DiffSensei显著优于现有模型,在角色可定制性方面实现重要进展。

原文摘要 · Abstract (English)

Story visualization, the task of creating visual narratives from textual descriptions, has seen progress with text-to-image generation models. However, these models often lack effective control over character appearances and interactions, particularly in multi-character scenes. To address these limitations, we propose a new task: \textbf{customized manga generation} and introduce \textbf{DiffSensei}, an innovative framework specifically designed for generating manga with dynamic multi-character control. DiffSensei integrates a diffusion-based image generator with a multimodal large language model (MLLM) that acts as a text-compatible identity adapter. Our approach employs masked cross-attention to seamlessly incorporate character features, enabling precise layout control without direct pixel transfer. Additionally, the MLLM-based adapter adjusts character features to align with panel-specific text cues, allowing flexible adjustments in character expressions, poses, and actions. We also introduce \textbf{MangaZero}, a large-scale dataset tailored to this task, containing 43,264 manga pages and 427,147 annotated panels, supporting the visualization of varied character interactions and movements across sequential frames. Extensive experiments demonstrate that DiffSensei outperforms existing models, marking a significant advancement in manga generation by enabling text-adaptable character customization. The project page is https://jianzongwu.github.io/projects/diffsensei/.

漫画生成多模态扩散模型角色控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。