用自适应提示生成人格测评题,效果比传统方法更好。
Prompt Engineering for Scale Development in Generative Psychometrics
- 采用自适应提示策略生成测评题,减少冗余并提升质量。
- 自适应提示使生成题库的结构效度更高,尤其配合强模型时优势明显。
- 适合研究生成式心理测量、AI辅助测评系统的设计者参考。
本蒙特卡洛模拟研究了提示工程策略如何影响大语言模型(LLM)在生成式心理测量框架AI-GENIE中生成的人格测评题质量。针对五大性格特质的题库通过零样本、少样本、角色导向和自适应提示等多种设计,结合不同模型温度与LLM生成,再经网络心理测量方法评估与精简。所有条件下,AI-GENIE在精简后均显著提升结构效度,其增益程度与原始题库质量呈反比。提示设计对精简前后题库质量均有显著影响。自适应提示持续优于非自适应策略,显著降低语义冗余,提高预精简结构效度,并保留更大题库规模,尤其在搭配新型高容量模型时。该优势在多数模型中对温度设置不敏感,表明自适应提示可缓解创意与心理测量一致性之间的权衡。唯一例外是GPT-4o在高温下表现下降,显示其对自适应约束在高随机性下的敏感性。总体表明,自适应提示在此场景下最优,且其效益随模型能力增强而扩大,推动进一步探索模型-提示交互在生成式心理测量流程中的作用。
原文摘要 · Abstract (English)
This Monte Carlo simulation examines how prompt engineering strategies shape the quality of large language model (LLM)--generated personality assessment items within the AI-GENIE framework for generative psychometrics. Item pools targeting the Big Five traits were generated using multiple prompting designs (zero-shot, few-shot, persona-based, and adaptive), model temperatures, and LLMs, then evaluated and reduced using network psychometric methods. Across all conditions, AI-GENIE reliably improved structural validity following reduction, with the magnitude of its incremental contribution inversely related to the quality of the incoming item pool. Prompt design exerted a substantial influence on both pre- and post-reduction item quality. Adaptive prompting consistently outperformed non-adaptive strategies by sharply reducing semantic redundancy, elevating pre-reduction structural validity, and preserving substantially larger item pool, particularly when paired with newer, higher-capacity models. These gains were robust across temperature settings for most models, indicating that adaptive prompting mitigates common trade-offs between creativity and psychometric coherence. An exception was observed for the GPT-4o model at high temperatures, suggesting model-specific sensitivity to adaptive constraints at elevated stochasticity. Overall, the findings demonstrate that adaptive prompting is the strongest approach in this context, and that its benefits scale with model capability, motivating continued investigation of model--prompt interactions in generative psychometric pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。