构建跨科学与艺术的视觉模式数据集,连接图像与生成机制。
SciTextures: Collecting and Connecting Visual Patterns, Models, and Code Across Science and Art
- 用智能代理自动收集并标准化100,000张图像与1,270个生成模型
- 验证视觉语言模型能从真实图像反推并复现生成过程
- 适合研究跨模态理解、科学可视化与生成模型的学者
将视觉模式与其生成机制相联系,是深层次视觉理解的核心。云层、波浪、城市与森林生长、材料及地貌形成等,皆为底层机制产生的视觉模式。本文提出SciTextures数据集,涵盖科学、技术与艺术领域超过1,270种模型和10万张纹理与模式图像,来自物理、化学、生物、社会学、技术、数学及艺术。该数据集通过智能代理流程自动采集、实现并标准化科学与生成模型,并自主发明新方法生成视觉模式。它支持对视觉语言模型(VLM)在关联视觉模式与生成代码方面的能力进行系统评估,也能识别同一机制生成的不同模式。我们测试了模型根据真实世界现象的自然图像,推断并重建其生成机制的能力:输入图像后,要求模型输出可运行的代码以模拟生成对应图像,并与原图对比。结果表明,当前VLM可在多抽象层次上理解并模拟物理系统。数据集与代码已开源:https://zenodo.org/records/17485502
原文摘要 · Abstract (English)
The ability to connect visual patterns with the processes that form them represents one of the deepest forms of visual understanding. Textures of clouds and waves, the growth of cities and forests, or the formation of materials and landscapes are all examples of patterns emerging from underlying mechanisms. We present the SciTextures dataset, a large-scale collection of textures and visual patterns from all domains of science, tech, and art, along with the models and code that generate these images. Covering over 1,270 different models and 100,000 images of patterns and textures from physics, chemistry, biology, sociology, technology, mathematics, and art, this dataset offers a way to explore the deep connection between the visual patterns that shape our world and the mechanisms that produce them. Built through an agentic AI pipeline that autonomously collects, implements, and standardizes scientific and generative models. This AI pipeline is also used to autonomously invent and implement novel methods for generating visual patterns and textures. SciTextures enables systematic evaluation of vision language models (VLM's) ability to link visual patterns to the models and code that generate them, and to identify different patterns that emerge from the same underlying process. We also test VLMs ability to infer and recreate the mechanisms behind visual patterns by providing a natural image of a real-world phenomenon and asking the AI to identify and code a model of the process that formed it, then run this code to generate a simulated image that is compared to the reference image. These benchmarks reveal that VLM's can understand and simulate physical systems beyond visual patterns at multiple levels of abstraction. The dataset and code are available at: https://zenodo.org/records/17485502
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。