通过系统化提示工程提升大模型对心理构念的识别准确率
Empirical Prompt Engineering for Construct Identification with Large Language Models
- 组合多种定义、指令和示例生成提示,筛选最优组合
- 在心理构念识别上使模型与人类判断一致率显著提升
- 适合需要高可信度判别的心理学研究场景
由于架构设计和海量预训练数据,大语言模型(LLMs)在文本分类任务中表现强劲。然而,其分类结果对提示词表述极为敏感,尤其在心理学等具有潜在复杂理论驱动构念的领域。本文提出并评估了一种系统性提示工程框架,用于提升心理构念识别效果。通过组合随机选取的多种构念定义、任务指令、编码指导及示例生成提示,并在训练集上实证选择表现最佳的组合,显著提高了LLM与人类分类的一致性。相比之下,角色扮演、链式思维推理和解释等提示策略带来的改进较小且不一致。该结论在多个模型和构念上均成立。整体而言,本方法为关键场景下提升人机分类一致性提供了实用、系统且理论敏感的解决方案。
原文摘要 · Abstract (English)
Due to their architecture and vast pre-training data, large language models (LLMs) demonstrate strong performance on text classification tasks. However, LLM classifications are highly responsive to prompt wording, particularly, as we show, in domains like psychology, where constructs are often latent, complex, and theory driven. Here, we present and evaluate a systematic framework for improving psychological construct identification through prompt engineering. We combinatorially generate prompts by appending random selections of multiple variants of construct definitions, task instructions, coding guidance, and examples. Empirically selecting the highest performing of these combinations in a training dataset substantially improves alignment between LLM and human classifications. In contrast, prompting techniques such as personas, chain-of-thought reasoning, and explanations provide smaller and less consistent improvements. This finding holds across multiple models and constructs. Overall, the approach we describe offers a practical, systematic, and theory-aware method for increasing the alignment between human and LLM classifications in settings where validity is critical.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。