提出四维框架评估AI生成故事的创造力,突破传统参考依赖评价局限。
Evaluation Framework for AI Creativity: A Case Study Based on Story Generation
- 构建包含新颖性、价值、契合度与共鸣感的四维评估体系
- 115名读者参与实证发现创造性判断分阶段显现且反思后更一致
- 适用于需要深度理解创意质量的研究者与内容平台
评估AI文本生成的创造力仍具挑战,因现有基于参考的指标无法捕捉创造力的主观性。本文提出一个结构化的AI故事生成评估框架,包含四个维度(新颖性、价值、契合度、共鸣感)和十一个子维度。通过‘Spike Prompting’控制生成并开展115名读者的众包研究,分析不同创意维度如何影响即时与反思性的人类创造力判断。结果表明,创造力评估呈层级而非累加模式,各维度在判断不同阶段凸显,且反思性评估显著改变评分结果与评分者间一致性。这些发现支持本框架有效揭示参考式评价所掩盖的创造力维度。
原文摘要 · Abstract (English)
Evaluating creative text generation remains a challenge because existing reference-based metrics fail to capture the subjective nature of creativity. We propose a structured evaluation framework for AI story generation comprising four components (Novelty, Value, Adherence, and Resonance) and eleven sub-components. Using controlled story generation via ``Spike Prompting'' and a crowdsourced study of 115 readers, we examine how different creative components shape both immediate and reflective human creativity judgments. Our findings show that creativity is evaluated hierarchically rather than cumulatively, with different dimensions becoming salient at different stages of judgment, and that reflective evaluation substantially alters both ratings and inter-rater agreement. Together, these results support the effectiveness of our framework in revealing dimensions of creativity that are obscured by reference-based evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。