让AI画画更符合个人审美,用艺术构图原则精准控制生成效果。
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
- 用艺术家的构图原则定义视觉美学,实现个性化控制。
- 通过轻量适配器实现10种构图参数可控,提升生成图像美感。
- 构建大型艺术数据集CompArt,支持美学评估与训练。
文本到图像扩散模型虽能生成高质量图像,但常产出缺乏美感的结果,因其训练数据包含大量网络随机图像。现有方法假设美学是普适的,限制了个性化表达。本文提出美学对齐新任务,旨在将用户指定的审美偏好与生成结果对齐。受艺术作品启发,我们采用艺术家常用的构图原则(Principles of Art, PoA)来形式化视觉美学。为此,我们构建了CompArt——一个基于WikiArt的大规模构图分析数据集,由多模态大模型标注了PoA标签。利用大模型的表达能力并训练轻量级可迁移适配器,我们证明扩散模型可通过用户指定的PoA条件实现10种构图控制。此外,设计了合理的评估框架验证方法有效性。
原文摘要 · Abstract (English)
Text-to-Image (T2I) diffusion models (DM) have garnered widespread adoption due to their capability in generating high-fidelity outputs and accessibility to anyone able to put imagination into words. However, DMs are often predisposed to generate unappealing outputs, much like the random images on the internet they were trained on. Existing approaches to address this are founded on the implicit premise that visual aesthetics is universal, which is limiting. Aesthetics in the T2I context should be about personalization and we propose the novel task of aesthetics alignment which seeks to align user-specified aesthetics with the T2I generation output. Inspired by how artworks provide an invaluable perspective to approach aesthetics, we codify visual aesthetics using the compositional framework artists employ, known as the Principles of Art (PoA). To facilitate this study, we introduce CompArt, a large-scale compositional art dataset building on top of WikiArt with PoA analysis annotated by a capable Multimodal LLM. Leveraging the expressive power of LLMs and training a lightweight and transferrable adapter, we demonstrate that T2I DMs can effectively offer 10 compositional controls through user-specified PoA conditions. Additionally, we design an appropriate evaluation framework to assess the efficacy of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。