arXiv:2606.30561cs.AIcs.CV2026-06

为创意AI设计新评估基准,区分专业共识与审美差异。

The Human Creativity Benchmark

论文配图:The Human Creativity Benchmark
图 1 · 摘自论文原文
  • 用成对偏好和多维度评分分离创作中的共识与分歧
  • 发现技术正确性易达成共识,美学方向则因人而异
  • 适合关注AI创意能力评估的学者与产品设计者

当前AI评估框架将评价者分歧视为噪声,但在创意领域,专业意见分歧反映真实品味差异,而非测量误差。我们主张创意AI评估应保留两种信号:收敛性(专业人士对规范做法的一致认可)与发散性(个体品味的合理差异)。为此提出人类创造力基准(HCB),通过收集15,000份跨五个创意领域、三个工作阶段(构思、原型、优化)的专业判断,包括成对偏好、提示契合度、可用性、视觉吸引力的标量评分及定性理由。结果表明,收敛集中在可验证维度如技术正确性和视觉层次,发散则集中于审美方向和概念风险等品味驱动维度。无模型在所有阶段表现均优。将两种信号合并为单一质量指标会丢弃最有效的信息:哪些方面需模型准确,哪些应保持可引导性。

原文摘要 · Abstract (English)

Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflects genuine differences in taste, not measurement error. We argue that evaluating creative AI requires preserving two distinct signals: convergence, where professionals align around shared best practices, and divergence, where individual taste legitimately varies. We present the Human Creativity Benchmark (HCB), a benchmark that operationalizes this separation by collecting pairwise preferences, scalar ratings on prompt adherence, usability, and visual appeal, and qualitative rationale from domain professionals. Across 15,000 professional judgments spanning five creative domains and three workflow phases (ideation, mockup, refinement), we find that convergence concentrates on verifiable dimensions like technical correctness and visual hierarchy, while divergence concentrates on taste-driven dimensions like aesthetic direction and conceptual risk. No model excels uniformly across all phases. Collapsing these signals into a single quality metric discards the most actionable information: where models must be correct versus where they should remain steerable.

AI评估创意生成人类反馈多维评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。