用大模型模拟自闭症等神经多样性成人的心理特征,效果优于随机猜测。
Large Language Models as Simulative Agents for Neurodivergent Adult Psychometric Profiles
- 基于结构化访谈内容,让大模型扮演个体回答心理量表
- GPT-4o表现最佳,对注意力缺陷等特质模拟准确率显著高于随机
- 适合用于早期心理测量研究中的虚拟被试,但细节维度仍有局限
成人神经多样性(包括注意力缺陷多动障碍、高功能自闭症谱系障碍和认知脱离综合征)症状重叠严重,限制了标准心理测量工具的区分能力。尽管已有研究显示大语言模型(LLMs)可从定性数据生成人类心理测量响应,但其能否准确稳定地建模神经发育特征而非泛化人格特质仍不明确。本研究检验了在结构化开放访谈基础上,大模型是否能生成接近真实个体的心理测量响应,以及模拟结果是否对特质强度变化敏感。26名成年人完成29项开放式访谈及四项标准化自评量表(ASRS、BAARS-IV、AQ、RAADS-R)。使用GPT-4o与Qwen3-235B-A22B两个模型,根据访谈内容推断个体心理画像并以角色身份作答。通过群体比较、误差度量、完全匹配评分和随机基线对比评估准确性、可靠性和敏感性。两个模型在所有量表上均优于随机响应,其中GPT-4o表现更优,具有更高准确率与可重复性。模拟响应在ASRS、BAARS-IV与RAADS-R上与真实数据高度吻合,而AQ量表在注意细节等子维度存在特定局限。结果表明,基于访谈的大模型可生成连贯且显著优于随机水平的神经发育特征模拟,支持其作为早期心理测量研究中合成被试的潜力,同时揭示领域特异性约束。
原文摘要 · Abstract (English)
Adult neurodivergence, including Attention-Deficit/Hyperactivity Disorder (ADHD), high-functioning Autism Spectrum Disorder (ASD), and Cognitive Disengagement Syndrome (CDS), is marked by substantial symptom overlap that limits the discriminant sensitivity of standard psychometric instruments. While recent work suggests that Large Language Models (LLMs) can simulate human psychometric responses from qualitative data, it remains unclear whether they can accurately and stably model neurodevelopmental traits rather than broad personality characteristics. This study examines whether LLMs can generate psychometric responses that approximate those of real individuals when grounded in a structured qualitative interview, and whether such simulations are sensitive to variations in trait intensity. Twenty-six adults completed a 29-item open-ended interview and four standardized self-report measures (ASRS, BAARS-IV, AQ, RAADS-R). Two LLMs (GPT-4o and Qwen3-235B-A22B) were prompted to infer an individual psychological profile from interview content and then respond to each questionnaire in-role. Accuracy, reliability, and sensitivity were assessed using group-level comparisons, error metrics, exact-match scoring, and a randomized baseline. Both models outperformed random responses across instruments, with GPT-4o showing higher accuracy and reproducibility. Simulated responses closely matched human data for ASRS, BAARS-IV, and RAADS-R, while the AQ revealed subscale-specific limitations, particularly in Attention to Detail. Overall, the findings indicate that interview-grounded LLMs can produce coherent and above-chance simulations of neurodevelopmental traits, supporting their potential use as synthetic participants in early-stage psychometric research, while highlighting clear domain-specific constraints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。