arXiv:2510.08663cs.CLcs.AI2025-10

用大模型从文本中生成心理测评新题,提升测量精度。

Augmenting Rating-Scale Measures with Text-Derived Items Using the Information-Determined Scoring (IDS) Framework

  • 用大模型简单提示词给自由文本打分,生成可融入量表的新题目。
  • 实测显示,10题内精准度提升至46.3%,相当于多测6.3个原题。
  • 适合做心理评估优化、临床筛查或自适应测试的开发者参考。

心理测评常依赖评分量表,要求受访者将复杂体验压缩到预设类别中。尽管常伴随丰富但非结构化的文本数据,却很少用于测量目标特质,因缺乏与潜在量表的直接映射。本文提出信息确定评分(IDS)框架,利用大语言模型(LLMs)通过简单提示对自由文本响应打分,生成候选题目,与基线量表共同校准,并根据其对目标特质提供的心理测量信息保留。这区别于传统自动化文本评分,更注重信息增益而非对专家评分标准或人工标注数据的拟合。以抑郁为案例,在693名高中生及3,000人的匹配合成数据集上测试。在保留测试集中,将19项量表补充大模型生成题目后,测量精度和准确性显著提升,且与外部自杀风险指标的收敛效度更强。自适应模拟显示,大模型生成题目信息量等价于真实数据中增加6.3个、合成数据中增加16.0个原量表题。在10题后,46.3%的受试者达到标准误≤0.3,优于基线的35.5%(真实数据),合成数据中达60.4%对比34.7%。结果表明,IDS框架能有效利用未结构化文本增强现有心理测量,适用于临床健康等领域。

原文摘要 · Abstract (English)

Psychological assessments commonly rely on rating-scale items, which require respondents to condense complex experiences into predefined categories. Although rich, unstructured text is often captured alongside these scales, it rarely contributes to measuring the target trait because it lacks direct mapping to the latent scale. We introduce the Information-Determined Scoring (IDS) framework, where large language models (LLMs) score free-text responses with simple prompts to generate candidate items that are co-calibrated with a baseline scale and retained based on the psychometric information they provide about the target trait. This marks a conceptual departure from traditional automated text scoring by prioritising information gain over fidelity to expert rubrics or human-annotated data. Using depression as a case study, we developed and tested the method in upper-secondary students (n = 693) and a matched synthetic dataset (n = 3,000). Across held-out test sets, augmenting a 19-item rating-scale measure with LLM-derived items yielded significant improvements in measurement precision and accuracy, and stronger convergent validity with an external suicidality measure throughout the adaptive test. In adaptive simulations, LLM-derived items contributed information equivalent to adding up to 6.3 and 16.0 rating-scale items in real and synthetic data, respectively. This enabled earlier high-precision measurement: after 10 items, 46.3% of respondents reached SE <= .3 under the strongest augmented test compared with 35.5% at baseline in real data, and 60.4% versus 34.7% in synthetic data. These findings illustrate how the IDS framework leverages unstructured text to enhance existing psychological measures, with applications in clinical health and beyond.

心理测量大模型应用自适应测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。