arXiv:2409.00202cs.CL2024-09被引 5

用大模型自动生成可信的创造力测评题,提升测试有效性。

The creative psychometric item generator: a framework for item generation and validation using large language models

  • 基于心理测量学框架,迭代优化创意问题生成
  • 生成题目能显著激发应试者更高质量的创意回答
  • 适合教育评估与AI创造力测评场景

大型语言模型(LLMs)在需要高度创造力的工作流程中日益普及。尽管已有大量研究探讨了LLMs自身的创造力,但关于其能否为人类生成有效创造力测评题的研究仍十分有限,而创造力在现代经济中的作用正日益关键。本文提出一种基于心理测量学的框架——创意心理计量题生成器(CPIG),用于生成经典自由应答式创造力测试中的题目,即创造性问题解决(CPS)任务。CPIG结合多种基于LLM的生成器与评价器,通过迭代方式优化写作提示,使后续生成的题目能引发被试更富创造性的回应。实证结果显示,CPIG生成的题目具备良好的效度与信度,且该效果并非源于评估过程中的已知偏差。研究结果表明,利用LLM可自动构建对人类和人工智能均有效的创造力测评工具。

原文摘要 · Abstract (English)

Increasingly, large language models (LLMs) are being used to automate workplace processes requiring a high degree of creativity. While much prior work has examined the creativity of LLMs, there has been little research on whether they can generate valid creativity assessments for humans despite the increasingly central role of creativity in modern economies. We develop a psychometrically inspired framework for creating test items (questions) for a classic free-response creativity test: the creative problem-solving (CPS) task. Our framework, the creative psychometric item generator (CPIG), uses a mixture of LLM-based item generators and evaluators to iteratively develop new prompts for writing CPS items, such that items from later iterations will elicit more creative responses from test takers. We find strong empirical evidence that CPIG generates valid and reliable items and that this effect is not attributable to known biases in the evaluation process. Our findings have implications for employing LLMs to automatically generate valid and reliable creativity tests for humans and AI.

创造力测评大模型生成心理测量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。