arXiv:2602.06015cs.CL2026-02被引 1

大模型评估创伤后应激障碍严重程度,关键在上下文信息与推理策略。

A Systematic Evaluation of Large Language Models for PTSD Severity Estimation: The Role of Contextual Knowledge and Modeling Strategies

  • 通过提供详细症状定义和语境信息提升评估准确性
  • 增加推理步骤能显著提高预测精度,模型规模影响因开源与否而异
  • 集成监督模型与零样本大模型效果最佳,适合临床风险筛查

本研究基于1,437名个体的自然语言叙事和自评PTSD严重度数据,系统评估了11个前沿大语言模型在无监督(生成式)场景下的表现。通过改变提示中的上下文知识(如量表定义、分布概要、访谈问题)和建模策略(零样本/少样本、推理强度、模型规模、结构化子量表与直接评分、输出重标定、九种集成方法),发现:(a) 提供详细构念定义和叙述背景时,模型准确率甚至超过人类评定者与自评的一致性;(b) 推理投入越多,估计越准;(c) 开源模型(Llama、DeepSeek)在700亿参数后性能趋于平稳,闭源模型(gpt-o3-mini、gpt-5)随新版本持续提升;(d) 将监督模型与零样本大模型集成可获最优效果。此外,模型估计能区分PTSD与其他心理状态,并前瞻性预测未来心理健康支出。结果表明,上下文知识与建模策略显著影响大模型在PTSD评估中的准确性和临床价值。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being used in a zero-shot (generative) fashion to assess mental health conditions, yet we have limited knowledge on what factors affect their accuracy. In this study, we use a clinical dataset of natural language narratives and self-reported PTSD severity scores from 1,437 individuals to comprehensively evaluate the performance of 11 state-of-the-art LLMs. To understand the factors affecting model's assessment accuracy, we systematically varied (i) contextual knowledge prompted to the models like subscale definitions, distribution summary, and interview questions, and (ii) modeling strategies including zero-shot vs few shot, amount of reasoning effort, model sizes, structured subscales vs direct scalar prediction, output rescaling and nine ensemble methods. Our findings indicate that (a) LLMs are most accurate when provided with detailed construct definitions and context of the narrative, even exceeding human raters agreement with self-reported scores; (b) increased reasoning effort leads to better estimation accuracy; (c) performance of open-weight models (Llama, DeepSeek) plateaus beyond 70B parameters while closed-weight (gpt-o3-mini, gpt-5) alternatives improve with newer generations; and (d) best performance is achieved when ensembling a supervised model with the zero-shot LLMs. Beyond agreement with self-reports, LLMs' estimates discriminated PTSD severity from depression, anxiety, and alcohol use, and prospectively predicted future mental healthcare expenditure. Together, these results suggest that contextual knowledge and modeling strategies meaningfully affect accuracy and clinical utility of LLM-based assessments of PTSD severity.

PTSD评估大模型应用临床决策零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。