用检索增强生成艺术图像,让文字描述更精准地变成画作。
MythraGen: Two-Stage Retrieval Augmented Art Generation Framework

- 从艺术数据库中检索相似图像,用LoRA微调Stable Diffusion。
- 在WikiArt数据集上生成效果显著优于现有方法。
- 适合需要风格精准匹配的艺术创作场景。
文本到图像生成已取得快速发展,尤其得益于生成模型的进步。然而,在实现高质量、上下文准确的图像输出方面仍面临挑战,尤其是在艺术生成中忠实匹配文本描述。本文提出一种简单高效的检索增强生成框架MythraGen,通过将艺术检索机制与基于LoRA的模型微调相结合,用于文本到艺术图像生成。该方法从大规模艺术数据集中提取特征,通过融合艺术家特定风格与内容来优化生成过程。具体而言,从外部艺术数据库中检索与查询提示最相似的图像,并使用这些图像对Stable Diffusion进行LoRA微调,以生成期望的艺术作品。在WikiArt数据集上的实验结果和用户研究显示,所提方法能生成与用户输入高度匹配的艺术作品,显著优于现有方案。
原文摘要 · Abstract (English)
Text-to-image generation has seen rapid advancements, especially with the development of generative models. However, challenges remain in achieving high-quality, contextually accurate image outputs that faithfully match the provided textual descriptions, especially in artistic generation. In this paper, we present a simple yet efficient retrieval augmented generation framework, namely MythraGen, for text-to-artistic image generation by integrating an art retrieval mechanism with LoRA-based model fine-tuning. Our method extracts features from a large-scale art dataset, optimizing the generation process by combining artist-specific styles and content. Particularly, retrieved images from an external art database that have the highest similarity to the query prompt are used to finetune Stable Diffusion using LoRA for desired art generation. Experimental results and user studies on the WikiArt dataset show that our proposed method can generate artworks that closely match the user's input, significantly outperforming existing solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。