arXiv:2409.19946cs.CV2024-09被引 7

开源动漫图像生成模型Illustrious,实现超高清画质与精细还原。

Illustrious: an Open Advanced Illustration Model

论文配图:Illustrious: an Open Advanced Illustration Model
图 1 · 摘自论文原文
  • 通过控制批次大小和丢弃率,加速概念激活学习。
  • 训练分辨率提升至2000万像素以上,精准还原角色结构。
  • 采用多层级标签描述,提升生成内容个性化能力。

本文分享了在文本到图像动漫生成模型Illustrious中实现顶尖画质的洞察。为达成高分辨率、动态色彩范围及强还原能力,我们聚焦三项关键改进:首先,深入研究批次大小与丢弃率控制,加速基于可调控标记的概念激活学习;其次,提高训练图像分辨率,显著增强角色解剖结构的精确呈现,结合适当方法将生成能力拓展至2000万像素以上;最后,提出优化的多层级标题系统,涵盖所有标签与多样化自然语言描述,成为模型发展的关键因素。通过大量分析与实验,Illustrious在动画风格生成上表现卓越,优于广泛使用的主流模型,推动更便捷的定制化与个性化应用。我们计划陆续公开更新后的Illustrious模型系列,并制定持续优化方案。

原文摘要 · Abstract (English)

In this work, we share the insights for achieving state-of-the-art quality in our text-to-image anime image generative model, called Illustrious. To achieve high resolution, dynamic color range images, and high restoration ability, we focus on three critical approaches for model improvement. First, we delve into the significance of the batch size and dropout control, which enables faster learning of controllable token based concept activations. Second, we increase the training resolution of images, affecting the accurate depiction of character anatomy in much higher resolution, extending its generation capability over 20MP with proper methods. Finally, we propose the refined multi-level captions, covering all tags and various natural language captions as a critical factor for model development. Through extensive analysis and experiments, Illustrious demonstrates state-of-the-art performance in terms of animation style, outperforming widely-used models in illustration domains, propelling easier customization and personalization with nature of open source. We plan to publicly release updated Illustrious model series sequentially as well as sustainable plans for improvements.

动漫生成高分辨率开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。