用虚拟受访者模拟心理测评题,提升大模型测评的准确性。
Psychometric Item Validation Using Virtual Respondents with Trait-Response Mediators
- 用大模型生成人格特质的中介因素,模拟多样化答题行为
- 在三大理论框架下验证,显著提升题目与目标特质的相关性
- 适合想低成本开发心理测评题的研究者和开发者
随着心理测量问卷越来越多地用于评估大语言模型(LLMs)的人格特质,为大模型量身定制可扩展的题目生成需求日益增长。关键挑战在于确保生成题目的构念效度——即是否真正测量了预期特质。传统方法依赖昂贵的大规模人类数据收集。为提高效率,我们提出一种基于大模型的虚拟被试仿真框架。核心思路是考虑中介因素:同一特质可能通过不同中介导致差异化的答题反应。通过模拟具有多样中介的被试,我们识别出在各类中介下仍与目标特质强相关的题目。在五大性格理论(Big5)、施瓦茨价值观理论(Schwartz)和优势与美德理论(VIA)上的实验表明,我们的中介生成方法和仿真框架能有效识别高效度题目。大模型展现出从特质定义生成合理中介并模拟被试行为的能力。本研究提出的问题设定、评价指标、方法论及数据集,为低成本问卷开发和深入理解大模型如何模拟人类问卷响应开辟新路径。我们已开源数据集与代码,以支持后续研究。
原文摘要 · Abstract (English)
As psychometric surveys are increasingly used to assess the traits of large language models (LLMs), the need for scalable survey item generation suited for LLMs has also grown. A critical challenge here is ensuring the construct validity of generated items, i.e., whether they truly measure the intended trait. Traditionally, this requires costly, large-scale human data collection. To make it efficient, we present a framework for virtual respondent simulation using LLMs. Our central idea is to account for mediators: factors through which the same trait can give rise to varying responses to a survey item. By simulating respondents with diverse mediators, we identify survey items that yield responses robustly correlated with intended traits across these mediators. Experiments on three psychological trait theories (Big5, Schwartz, VIA) show that our mediator generation methods and simulation framework effectively identify high-validity items. LLMs demonstrate the ability to generate plausible mediators from trait definitions and to simulate respondent behavior for item validation. Our problem formulation, metrics, methodology, and dataset open a new direction for cost-efficient survey development and a deeper understanding of how LLMs simulate human survey responses. We release our dataset and code to support future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。