用大模型让普通人也能轻松创作动态视觉效果。
AI Co-Artist: A LLM-Powered Framework for Interactive GLSL Shader Animation Evolution
- 通过大模型理解用户意图,自动生成和演化GLSL着色器代码。
- 用户无需编程即可通过直观操作生成专业级实时动画效果。
- 适合艺术创作者、设计新手及希望快速原型化视觉作品的团队。
创意编程与实时着色器开发是互动数字艺术的前沿,使艺术家、设计师和爱好者能够基于声音或用户交互等实时输入生成令人惊叹的复杂视觉效果。然而,尽管GLSL工具潜力巨大,其陡峭的学习曲线和对编程能力的要求,仍对初学者乃至非技术背景的资深艺术家构成重大障碍。本文提出AI Co-Artist,一个基于大语言模型(如GPT-4)的交互式系统,通过用户友好的可视化界面支持用户以迭代方式演化和优化GLSL着色器。该系统借鉴Picbreeder平台的用户引导演化理念,使用户无需编写或理解代码,即可通过直观交互逐步进化出视觉艺术作品。作为创意伙伴与技术助手,系统帮助用户探索广阔的实时视觉艺术生成空间。通过结构化用户研究与定性反馈评估,我们证明了该系统显著降低了着色器创作的技术门槛,提升了创作质量,并支持广泛用户群体产出专业级视觉效果。此外,我们认为这一范式具有广泛可迁移性:结合大模型在语义理解与程序合成方面的双重优势,该方法可推广至网站布局生成、建筑可视化、产品原型设计和信息图等领域。
原文摘要 · Abstract (English)
Creative coding and real-time shader programming are at the forefront of interactive digital art, enabling artists, designers, and enthusiasts to produce mesmerizing, complex visual effects that respond to real-time stimuli such as sound or user interaction. However, despite the rich potential of tools like GLSL, the steep learning curve and requirement for programming fluency pose substantial barriers for newcomers and even experienced artists who may not have a technical background. In this paper, we present AI Co-Artist, a novel interactive system that harnesses the capabilities of large language models (LLMs), specifically GPT-4, to support the iterative evolution and refinement of GLSL shaders through a user-friendly, visually-driven interface. Drawing inspiration from the user-guided evolutionary principles pioneered by the Picbreeder platform, our system empowers users to evolve shader art using intuitive interactions, without needing to write or understand code. AI Co-Artist serves as both a creative companion and a technical assistant, allowing users to explore a vast generative design space of real-time visual art. Through comprehensive evaluations, including structured user studies and qualitative feedback, we demonstrate that AI Co-Artist significantly reduces the technical threshold for shader creation, enhances creative outcomes, and supports a wide range of users in producing professional-quality visual effects. Furthermore, we argue that this paradigm is broadly generalizable. By leveraging the dual strengths of LLMs-semantic understanding and program synthesis, our method can be applied to diverse creative domains, including website layout generation, architectural visualizations, product prototyping, and infographics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。