用问答方式筛选图像,让AI生成更精准的主题图像。
Selecting Fine-Tuning Examples by Quizzing VLMs
- 用视觉模型问答打分,自动选高质量训练图像。
- 用少量精选图像即可生成逼真且主题一致的图像。
- 适合需要精准控制生成内容的研究者和创作者。
在为特定主题微调文生图扩散模型时,选择优质训练样本是关键挑战。从维基百科公共图库等质量参差的图像集合中微调,常导致生成效果不佳。然而,能准确体现目标概念的图像(如女性蓝山雀)有助于生成具有典型特征(如蓝色翅膀、灰色胸羽)的图像。本文提出QZLoRA框架,通过QuizRank方法,将图像视为教育干预,对视觉语言模型(VLM)进行‘提问’以自动排序并筛选图像,用于低秩适应(LoRA)微调。实验表明,QZLoRA仅需较少样本即可生成更对齐、更逼真的图像;且微调后的模型还能生成风格化但同样具代表性的插画。结果表明,自动化视觉推理与参数高效微调结合,在主题自适应生成建模中具有巨大潜力。
原文摘要 · Abstract (English)
A challenge in fine-tuning text-to-image diffusion models for specific topics is to select good examples. Fine-tuning from image sets of varying quality, such as Wikipedia Commons, will often produce poor output. However, training images that \textit{do} exemplify the target concept (e.g., a \textit{female Mountain Bluebird}) help ensure that the generated images are similarly representative (e.g., have the prototypical blue-wings and gray chest). In this work, we propose QZLoRA, a framework to select images for low-rank adaptation (LoRA). The approach leverages QuizRank, a method to automatically rank images by treating them as an `educational intervention' and `quizzing' a VLM. We demonstrate that QZLoRA can produce better aligned, photorealistic images with fewer samples. We also show that these fine-tuned models can produce stylized that are similarly representative (i.e., illustrations). Our results highlight the promise of combining automated visual reasoning with parameter-efficient fine-tuning for topic-adaptive generative modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。