用视觉语言模型引导人工生命演化,实现复杂目标的自动进化。
Guiding Evolution of Artificial Life Using Vision-Language Models
- 用第二个基础模型根据视觉历史生成新演化目标,驱动开放性演化。
- 单目标演化提升视觉新颖性,序列目标演化增强演化过程连贯性。
- 适合对生成式模型与人工生命交叉研究感兴趣的学者参考。
基础模型(FMs)为人工生命(ALife)领域带来了新突破,提供了自动化搜索ALife模拟的强大工具。以往工作利用视觉语言模型(VLMs)将ALife模拟与自然语言目标对齐。我们基于自动人工生命搜索(ASAL),提出ASAL++,一种由多模态基础模型引导的开放式搜索方法。采用第二个基础模型,基于模拟的视觉历史生成新的演化目标,从而推动演化轨迹向更复杂的目標发展。我们探索两种策略:(1) 在每轮迭代中让模拟匹配单一新提示(演化监督目标:EST);(2) 让模拟匹配生成的整个提示序列(演化时间目标:ETT)。我们在Lenia底座上使用Gemma-3进行实验,结果表明,EST促进更高视觉新颖性,而ETT则带来更连贯、可解释的演化序列。结果表明,ASAL++为以基础模型驱动的、具有开放性的ALife发现指明了新方向。
原文摘要 · Abstract (English)
Foundation models (FMs) have recently opened up new frontiers in the field of artificial life (ALife) by providing powerful tools to automate search through ALife simulations. Previous work aligns ALife simulations with natural language target prompts using vision-language models (VLMs). We build on Automated Search for Artificial Life (ASAL) by introducing ASAL++, a method for open-ended-like search guided by multimodal FMs. We use a second FM to propose new evolutionary targets based on a simulation's visual history. This induces an evolutionary trajectory with increasingly complex targets. We explore two strategies: (1) evolving a simulation to match a single new prompt at each iteration (Evolved Supervised Targets: EST) and (2) evolving a simulation to match the entire sequence of generated prompts (Evolved Temporal Targets: ETT). We test our method empirically in the Lenia substrate using Gemma-3 to propose evolutionary targets, and show that EST promotes greater visual novelty, while ETT fosters more coherent and interpretable evolutionary sequences. Our results suggest that ASAL++ points towards new directions for FM-driven ALife discovery with open-ended characteristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。