用大视觉语言模型复现开放生成系统,探索AI能否像人类一样持续创造新内容。
In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models

- 用前沿视觉语言模型替代人类,模拟图片进化过程
- 生成图像在语义新颖性和视觉复杂度上低于人类历史成果
- 引入噪声和记忆机制后,创造性表现略有提升,适合研究开放生成
当前学术与产业界正大力推动人工智能助手自动化科学、技术和创造性生产。历史上,这些过程的人类形态具有核心特征——开放性:能够持续生成看似无穷无尽的新颖且有意义的形态。人工代理是否具备此类富有成效的无指导发现能力?为此,我们以经典的人类驱动开放搜索系统Picbreeder为案例,其通过用户协作交互演化小型神经网络生成多样化图像。本文用前沿视觉语言模型(VLMs)替代人类用户进行复现。观察到系统输出与历史人类基准存在明显定性差异,并通过谱系复杂性、视觉与语义显著性及新颖性等指标进行表征。为进一步识别导致差异的因果因素,我们考察了在代理选择过程中加入探索性噪声、代理间行为多样性以及基于过往行为记忆的叙事推进机制的影响。代码已开源:https://github.com/smearle/picbreeder-vlm。
原文摘要 · Abstract (English)
We are in the midst of large-scale industrial and academic efforts to automate the processes of scientific, technological and creative production through AI-driven assistants. Historically, a fundamental property of these processes in their human form has been their open-endedness: their capacity for generating a seemingly endless supply of novel and meaningful new forms. Do artificial agents have any capacity for such fruitful unguided discovery? To answer this question, we turn to Picbreeder, the canonical exemplar of human-driven open-ended search, in which users collaboratively generated a diverse library of images through interactive evolution of small neural networks. We replicate Picbreeder, replacing human users with frontier Vision Language Models (VLMs). We observe clear qualitative differences between the output of our system and the historical human baseline, and attempt to characterize them using metrics of phylogenetic complexity and visual and semantic salience and novelty. In an effort to identify some of the causal factors contributing these differences, we study the addition of exploratory noise to the agents' selection process, of behavioral diversity between agents, and of narrative momentum in the form of memory of past actions. We make our code available at https://github.com/smearle/picbreeder-vlm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。