arXiv:2509.09702cs.CLcs.AI2025-09被引 2

评测大模型在品牌营销创意上的表现,发现无模型全面领先。

Creativity Benchmark: A benchmark for marketing creativity for large language models

  • 构建涵盖100个品牌的创意评测基准,覆盖三类提示。
  • 顶级模型仅比最低模型胜率高61%,无模型全程领先。
  • 强调人工评估必要性,自动化评分不可替代人类判断。

我们提出Creativity Benchmark,用于评估大语言模型在营销创意方面的表现。该基准包含100个品牌(12个类别)和三种提示类型(洞察、创意、天马行空)。基于678名从业创意人员对11,012次匿名对比的偏好数据,通过Bradley-Terry模型分析显示,模型表现高度集中,无模型在所有品牌或提示类型上占据绝对优势:顶尖与垫底模型之间的差距Δθ≈0.45,对应一对一胜率约61%。我们还利用余弦距离分析模型多样性,考察其对提示重构的敏感性。对比三种LLM作为评判者与人工评分的结果,发现相关性弱且不一致,存在显著的评判偏差,表明自动化评判无法取代人工评估。传统创意测试在品牌约束任务中也仅部分适用。总体结果强调需依赖专家人工评估,并建立注重多样性的工作流程。

原文摘要 · Abstract (English)

We introduce Creativity Benchmark, an evaluation framework for large language models (LLMs) in marketing creativity. The benchmark covers 100 brands (12 categories) and three prompt types (Insights, Ideas, Wild Ideas). Human pairwise preferences from 678 practising creatives over 11,012 anonymised comparisons, analysed with Bradley-Terry models, show tightly clustered performance with no model dominating across brands or prompt types: the top-bottom spread is $Δθ\approx 0.45$, which implies a head-to-head win probability of $0.61$; the highest-rated model beats the lowest only about $61\%$ of the time. We also analyse model diversity using cosine distances to capture intra- and inter-model variation and sensitivity to prompt reframing. Comparing three LLM-as-judge setups with human rankings reveals weak, inconsistent correlations and judge-specific biases, underscoring that automated judges cannot substitute for human evaluation. Conventional creativity tests also transfer only partially to brand-constrained tasks. Overall, the results highlight the need for expert human evaluation and diversity-aware workflows.

大模型评测营销创意人类评估多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。