用视觉语言模型把图片变代码,自动生成可模拟的图像。
Coding the Visual World: From Image to Simulation Using Vision Language Models
- 给模型看图,让它写代码生成对应图像,实现从图像到模拟的转化。
- 顶尖模型能理解复杂系统的多层抽象,但难还原细节和局部排列。
- 适合对生成式建模、跨域系统理解感兴趣的科研人员。
构建对世界的心理模型是理解的核心。同样,视觉理解可视为对图像所呈现系统建立代表性模型的能力。本文通过Im2Sim方法,探索视觉语言模型(VLM)识别并模拟图像中系统与机制的能力。将真实世界系统(如城市、云、植被)的自然图像输入VLM,要求其描述系统并编写可执行的代码以模拟生成该图像。生成的合成图像与原图对比评估。实验涵盖物理系统(波、光、云)、植被、城市、材料及地质构造等复杂涌现系统。分析表明,领先VLM(GPT、Gemini)具备在多层抽象和广泛领域内理解复杂多组件系统的能力;但对图像中的精细细节和低层级模式排列再现能力有限。结果揭示出一种有趣不对称:模型兼具高层次深度视觉理解,却在细粒度感知上表现不足。
原文摘要 · Abstract (English)
The ability to construct mental models of the world is a central aspect of understanding. Similarly, visual understanding can be viewed as the ability to construct a representative model of the system depicted in an image. This work explores the capacity of Vision Language Models (VLMs) to recognize and simulate the systems and mechanisms depicted in images using the Im2Sim methodology. The VLM is given a natural image of a real-world system (e.g., cities, clouds, vegetation) and is tasked with describing the system and writing code that simulates and generates it. This generative code is then executed to produce a synthetic image, which is compared against the original. This approach is tested on various complex emergent systems, ranging from physical systems (waves, lights, clouds) to vegetation, cities, materials, and geological formations. Through analysis of the models and images generated by the VLMs, we examine their understanding of the systems in images. The results show that leading VLMs (GPT, Gemini) have the ability to understand and model complex, multi-component systems across multiple layers of abstraction and a wide range of domains. At the same time, the VLMs exhibit limited ability to replicate fine details and low-level arrangements of patterns in the image. These findings reveal an interesting asymmetry: VLMs combine high-level, deep visual understanding of images with limited perception of fine details.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。