LoRA权重可替代图像描述艺术风格,提升模型检索精度。
A LoRA is Worth a Thousand Pictures
- 用LoRA权重直接表征艺术风格,无需生成图像或原始训练数据。
- 在风格聚类任务中,LoRA表现优于CLIP、DINO等预训练特征。
- 适用于无训练图像信息的现实场景,支持零样本微调与模型溯源。
扩散模型与参数高效微调(PEFT)的进步使文本到图像生成与定制化变得广泛可用,其中低秩适配(LoRA)仅需少量数据与计算即可复现艺术家风格或主题。本文研究了LoRA权重与艺术风格的关系,证明仅凭LoRA权重即可有效描述风格,无需生成图像或知晓原始训练集。实验表明,基于LoRA的风格聚类性能优于传统预训练特征(如CLIP、DINO),其嵌入结构在定性与定量上均与图像嵌入高度相似。我们识别出多种定制模型检索场景,证实该方法在缺乏训练图像信息且需额外生成时仍能实现更精准检索。最后讨论了未来应用,如零样本LoRA微调与模型归属分析。
原文摘要 · Abstract (English)
Recent advances in diffusion models and parameter-efficient fine-tuning (PEFT) have made text-to-image generation and customization widely accessible, with Low Rank Adaptation (LoRA) able to replicate an artist's style or subject using minimal data and computation. In this paper, we examine the relationship between LoRA weights and artistic styles, demonstrating that LoRA weights alone can serve as an effective descriptor of style, without the need for additional image generation or knowledge of the original training set. Our findings show that LoRA weights yield better performance in clustering of artistic styles compared to traditional pre-trained features, such as CLIP and DINO, with strong structural similarities between LoRA-based and conventional image-based embeddings observed both qualitatively and quantitatively. We identify various retrieval scenarios for the growing collection of customized models and show that our approach enables more accurate retrieval in real-world settings where knowledge of the training images is unavailable and additional generation is required. We conclude with a discussion on potential future applications, such as zero-shot LoRA fine-tuning and model attribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。