用视觉语言模型评估设计美学,让算法自动变美。
Evolving to the Aesthetics of a Vision-Language Model

- 用CLIP-IQA或对比打分预测设计美感。
- 通过用户自定义提示进行两两比较,得排名。
- 艺术家实测验证,适合创意设计优化。
进化系统在生成字体、设计和音乐等创造性领域已取得显著成果,但如何设计能有效捕捉抽象输出美学的适应度函数仍是开放问题。本文探索两种基于视觉-语言模型(VLMs)评估种群美学的方法:第一种使用CLIP-IQA为每个设计预测美学分数;第二种则采用用户自定义提示,让候选设计两两竞争,胜者由VLM判定,再通过Glicko评分系统估算整体排名。我们在一个定制生成系统中开展案例研究,将所得排名与艺术家审美判断及其他评价方法结果对比,并记录艺术家使用这些方法演化设计的实际体验,批判性分析两种方法的优劣。
原文摘要 · Abstract (English)
Evolutionary systems have demonstrated remarkable results in creative domains, with recent applications in generative typography, design, and music. However, an open problem remains in designing fitness functions that effectively capture the desired aesthetics of abstract outputs. In this work, we explore two methods for evaluating the aesthetics of a population using Vision-Language Models (VLMs). The first method uses CLIP-IQA to predict an aesthetic score for each design. The second method instead pits candidates against each other, with winners determined by a VLM using a custom prompt specified by the user. The outcomes of these pairwise comparisons are then used to estimate a population ranking via the Glicko rating system. We present these methods in the context of a case study using a custom generative system and compare the resulting rankings with an artist's aesthetic ranking and those produced by other aesthetic evaluation techniques. Additionally, we document the artist's experience using these approaches to evolve designs, critically analysing the strengths and weaknesses of both methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。