arXiv:2508.10226cs.CLcs.AI2025-08被引 1

用大模型从访谈记录预测精神分裂症高危患者症状严重度,效果接近专业医生。

Using Large Language Models to Measure Symptom Severity in Patients At Risk for Schizophrenia

  • 用大语言模型直接分析未结构化访谈文本,预测BPRS评分。
  • 零样本预测与真实评估相关性达0.84(中位数),接近医生一致性。
  • 可跨语言应用,支持长期追踪,适合临床和研究场景。

精神分裂症临床高危(CHR)患者需密切监测症状以指导治疗。简明精神病评定量表(BPRS)是测量精神分裂症及其他精神病性障碍症状的可靠研究工具,但因需进行长时间结构化访谈,未广泛应用于临床。本研究利用大语言模型(LLMs),基于加速药物研发计划精神分裂症(AMP-SCZ)队列中409名CHR患者的临床访谈记录,预测其BPRS评分。尽管访谈未专门设计用于评估BPRS,LLM的零样本预测性能(中位数一致性:0.84,组内相关系数ICC:0.73)已接近人类评价者间的可靠性。进一步研究表明,LLM在跨语言评估(中位数一致性:0.88,ICC:0.70)及整合纵向信息的一次或少样本学习中具有显著潜力,有望提升并标准化对高危患者的评估。

原文摘要 · Abstract (English)

Patients who are at clinical high risk (CHR) for schizophrenia need close monitoring of their symptoms to inform appropriate treatments. The Brief Psychiatric Rating Scale (BPRS) is a validated, commonly used research tool for measuring symptoms in patients with schizophrenia and other psychotic disorders; however, it is not commonly used in clinical practice as it requires a lengthy structured interview. Here, we utilize large language models (LLMs) to predict BPRS scores from clinical interview transcripts in 409 CHR patients from the Accelerating Medicines Partnership Schizophrenia (AMP-SCZ) cohort. Despite the interviews not being specifically structured to measure the BPRS, the zero-shot performance of the LLM predictions compared to the true assessment (median concordance: 0.84, ICC: 0.73) approaches human inter- and intra-rater reliability. We further demonstrate that LLMs have substantial potential to improve and standardize the assessment of CHR patients via their accuracy in assessing the BPRS in foreign languages (median concordance: 0.88, ICC: 0.70), and integrating longitudinal information in a one-shot or few-shot learning approach.

精神分裂症大模型临床评估BPRS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。