构建可解释的视觉创造力模型,支持人机协同创作。
TraitSpaces: Towards Interpretable Visual Creativity for Human-AI Co-Creation
- 基于心理学与艺术家访谈,定义12种创造力特质。
- 部分特质在CLIP嵌入中预测可靠(R²≈0.64-0.68)。
- 提供可视化特质空间,助力有方向的人机共创。
我们提出一个基于心理学和艺术家经验的框架,用于建模四种领域中的视觉创造力:内在世界、外在世界、想象世界与道德世界。结合对实践艺术家的访谈与心理学理论,定义了12种涵盖情感、象征、文化与伦理维度的创造力特质。利用来自SemArt数据集的2万张艺术作品,通过GPT-4.1进行详尽、理论对齐的标注,评估这些特质从CLIP图像嵌入中学习的可行性。结果显示,如环境对话性(Environmental Dialogicity)和救赎弧线(Redemptive Arc)等特质预测可靠性较高(R²≈0.64–0.68),而记忆印记(Memory Imprint)等仍具挑战,凸显纯视觉编码的局限。除技术指标外,我们可视化“创造力特质空间”,展示如何沿红十字轴滑动以探索困境与重生主题的作品。本工作旨在将文化美学洞见与计算建模结合,不将创造力简化为数字,而是为艺术家、研究者与AI系统提供共享语言与可解释工具,实现有意义的协同创作。
原文摘要 · Abstract (English)
We introduce a psychologically grounded and artist-informed framework for modeling visual creativity across four domains: Inner, Outer, Imaginative, and Moral Worlds. Drawing on interviews with practicing artists and theories from psychology, we define 12 traits that capture affective, symbolic, cultural, and ethical dimensions of creativity.Using 20k artworks from the SemArt dataset, we annotate images with GPT 4.1 using detailed, theory-aligned prompts, and evaluate the learnability of these traits from CLIP image embeddings. Traits such as Environmental Dialogicity and Redemptive Arc are predicted with high reliability ($R^2 \approx 0.64 - 0.68$), while others like Memory Imprint remain challenging, highlighting the limits of purely visual encoding. Beyond technical metrics, we visualize a "creativity trait-space" and illustrate how it can support interpretable, trait-aware co-creation - e.g., sliding along a Redemptive Arc axis to explore works of adversity and renewal. By linking cultural-aesthetic insights with computational modeling, our work aims not to reduce creativity to numbers, but to offer shared language and interpretable tools for artists, researchers, and AI systems to collaborate meaningfully.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。