发现语言模型心理测评中提示词会引入显著干扰,影响结果可信度
The Unsampled Truth: Quantifying Prompt Artifacts in LM Psychometrics
- 用交叉实验分离提示词中的语义与非语义因素影响
- 非语义提示变化导致的响应差异超50%且超过人格设定本身
- 为心理测评类研究提供可量化的提示干扰检测方法
在使用语言模型进行心理测评时,研究者通常假设回应反映的是注入的人格设定和题目含义。本文通过诊断设计,将五个语义不同的基础人格设定与四个提示组件(人格描述、任务指令、题目表述、选项符号)的五种语义等价变体交叉组合,测量响应分布间的1-Wasserstein距离,并分解各组件对变异的贡献。该框架应用于13个开源小规模语言模型(0.6B至14B),测试了大五人格量表和短版黑暗三联征量表。结果显示,在多数模型中,任务指令与选项符号对响应分布的影响大于改写人格描述或题目本身;大量题目中,提示干扰解释的方差占比超过50%,非语义提示变化导致的响应变异甚至超过基础人格设定的影响。该框架使研究人员可在解读心理测评结果前量化提示干扰。
原文摘要 · Abstract (English)
When prompting language models for psychometric assessment, researchers assume that the responses reflect the injected persona and the meaning of the survey item. We test this premise using a diagnostic design that crosses five semantically distinct baseline personas with five semantically equivalent variants of each of four prompt components (persona wording, task instruction, item wording, option symbol). Measuring the 1-Wasserstein distance between the resulting response distributions and partitioning the variation among the five components allows for the separation of target effects from prompt artifacts. We apply the framework to 13 open-weight small language models (0.6B to 14B) on the Big Five Inventory and the Short Dark Triad. We find that in most models, the task instruction and option symbol displace response distributions further than paraphrasing the persona description or the item itself. For a substantial share of items, the artifact share of explained variation exceeds 50%; non-semantic changes of the prompt account for more response variation than the baseline personas. Our framework lets researchers quantify these prompt artifacts before interpreting psychometric output.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。