用大模型自动生成性格测评题,效率更高且质量可靠。
Automatic Item Generation for Personality Situational Judgment Tests with Large Language Models
- 设计优化提示词与温度参数,提升生成题目质量
- 在GPT-4和ChatGPT-5上稳定生成高质量题目
- 适合人力资源、心理测评等需高效建题的场景
通过情景判断测试(SJTs)进行人格评估相比传统自评量表具有独特优势,但其开发过程仍高度依赖专家,耗时费力。近年来大语言模型(LLMs)在自动题目生成(AIG)方面展现出潜力。本研究构建并评估了一种结构化、可推广的自动化人格SJTs生成框架,以GPT-4和ChatGPT-5为实证案例。研究1系统比较了提示词设计与温度设置对生成题目内容效度的影响,发现优化提示词与温度1.0在GPT-4上实现创意与准确性的最佳平衡。研究2通过多轮实验检验该方法在跨模型间的可复现性与泛化能力,结果表明在ChatGPT-5上仍能持续生成高质量题目。研究3评估了覆盖大五人格五个维度的生成题目的心理测量特性,结果显示多数维度具有令人满意的信度与效度,但在遵从性维度的收敛效度及部分准则效度方面存在局限。研究证明,该基于大模型的AIG方法可高效生成文化适切、心理测量学性能良好的人格测评题,效率优于或媲美传统方法。
原文摘要 · Abstract (English)
Personality assessment through situational judgment tests (SJTs) offers unique advantages over traditional Likert-type self-report scales, yet their development remains labor-intensive, time-consuming, and heavily dependent on subject matter experts. Recent advances in large language models (LLMs) have shown promise for automatic item generation (AIG). Building on these developments, the present study focuses on developing and evaluating a structured and generalizable framework for automatically generating personality SJTs, using GPT-4 and ChatGPT-5 as empirical examples. Three studies were conducted. Study 1 systematically compared the effects of prompt design and temperature settings on the content validity of LLM-generated items to develop an effective and stable LLM-based AIG approach for personality SJT. Results showed that optimized prompts and a temperature of 1.0 achieved the best balance of creativity and accuracy on GPT-4. Study 2 examined the cross-model generalizability and reproducibility of this automated SJT generation approach through multiple rounds. The results showed that the approach consistently produced reproducible and high-quality items on ChatGPT-5. Study 3 evaluated the psychometric properties of LLM-generated SJTs covering five facets of the Big Five personality traits. Results demonstrated satisfactory reliability and validity across most facets, though limitations were observed in the convergent validity of the compliance facet and certain aspects of criterion-related validity. These findings provide robust evidence that the proposed LLM-based AIG approach can produce culturally appropriate and psychometrically sound SJTs with efficiency comparable to or exceeding traditional methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。