arXiv:2608.06967cs.CL2026-08

测试大模型能否在无图像情况下生成原创视觉构思。

Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs

论文配图:Can Language Models Imagine Without Seeing? Ekphrasis: Measuring Visual Creative Ideation in Text-Only LLMs
图 1 · 摘自论文原文
  • 提出视觉创意构思评测框架Ekphrasis,覆盖抽象、组合、变换、改编四类任务。
  • 14个模型中,优秀表现者在有用性、表达力、新颖性上各有侧重,非仅靠流畅度。
  • 结果通过图像还原和盲评验证,证明文本级创意可独立于语言质量存在。

现有评估无法区分纯文本语言模型是否能在图像生成前产生视觉概念。流利的视觉描述可能掩盖视觉构想失败:答案看似有创意,实则重复常见视觉套路或无法生成具体场景。本文定义视觉创意构思(VCI)为生成有用、富有表现力且群体新颖的文本视觉方案的能力,并提出包含400个任务的Ekphrasis基准,涵盖抽象、组合、变换、适应四类。该基准采用匿名成对比较,结合维度专项清单,用布拉德利-特里模型聚合偏好,并利用类型化构思图将任务特定的群体陈词滥调转化为新颖性参照。在14个语言模型上的实验表明,VCI能有效分离有用性、表达力与新颖性,而非简单归结为语言流畅度;强模型虽总分相近,但三维度表现各异,且有用方案仍可能为视觉陈规。跨模态接地研究进一步显示,文本级VCI排序在忠实图像还原与盲图像偏好判断后仍保持稳定,支持Ekphrasis作为超越语言质量的视觉构想测量工具。

原文摘要 · Abstract (English)

Current evaluations do not isolate whether text-only language models can originate visual concepts before image generation. Fluent visual prose can hide visual-plan failures: an answer may appear creative while repeating familiar visual clichés or failing to specify a renderable scene. We define Visual Creative Ideation (VCI) as the ability to produce textual visual plans that are useful, expressive, and population-novel, and introduce Ekphrasis, a 400-task benchmark spanning Abstraction, Combination, Transformation, and Adaptation. Ekphrasis scores anonymized pairwise comparisons with dimension-specific checklists, aggregates preferences with Bradley-Terry models, and uses Typed Idea Graphs to convert task-specific population clichés into novelty references. Across 14 language models, VCI separates usefulness, expressiveness, and novelty rather than reducing to fluency: strong models achieve similar overall scores through different profiles, and useful plans can remain visually clichéd. A cross-modal grounding study further shows that text-level VCI ordering largely survives faithful rendering and blind image-level preference judgment, supporting Ekphrasis as a measure of visual ideation beyond prose quality.

视觉创意语言模型评测基准文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。