用大模型同时识别多种自杀风险因素,提升早期干预精准度。
Multi-Label Classification with Generative AI Models in Healthcare: A Case Study of Suicidality and Risk Factors
- 用生成式大模型直接完成多标签分类,避免传统二分类忽略共现风险。
- 微调后的GPT-3.5在部分匹配上达94%准确率,F1值0.91,表现优异。
- 发现模型易混淆自伤与自杀意图,适合临床研究与精神健康系统开发。
自杀仍是全球重大公共卫生危机,每年导致超过72万例死亡,另有数百万人受自杀意念(SI)和自杀未遂(SA)影响。早期识别包括SI、SA、接触自杀(ES)和非自杀性自伤(NSSI)在内的自杀相关因素(SrFs)对及时干预至关重要。以往研究多将自杀风险视为二分类任务,忽视多种风险因素共现的复杂性。本研究探索使用生成式大语言模型(如GPT-3.5和GPT-4.5)从精神科电子病历(EHRs)中进行多标签分类(MLC)。提出端到端的生成式多标签分类流程,并引入标签集级别评估指标及多标签混淆矩阵用于错误分析。微调后的GPT-3.5达到0.94的局部匹配准确率和0.91的F1分数;而经引导提示的GPT-4.5在各类标签集(包括罕见或少数类)上表现更优,体现更强的平衡性与鲁棒性。研究揭示系统性错误模式,如将SI与SA混淆,且模型倾向谨慎地过度标注。本工作不仅证明了生成式AI处理复杂临床分类任务的可行性,也为结构化非结构化病历数据以支持大规模临床研究与循证医学提供了范式。
原文摘要 · Abstract (English)
Suicide remains a pressing global health crisis, with over 720,000 deaths annually and millions more affected by suicide ideation (SI) and suicide attempts (SA). Early identification of suicidality-related factors (SrFs), including SI, SA, exposure to suicide (ES), and non-suicidal self-injury (NSSI), is critical for timely intervention. While prior studies have applied AI to detect SrFs in clinical notes, most treat suicidality as a binary classification task, overlooking the complexity of cooccurring risk factors. This study explores the use of generative large language models (LLMs), specifically GPT-3.5 and GPT-4.5, for multi-label classification (MLC) of SrFs from psychiatric electronic health records (EHRs). We present a novel end to end generative MLC pipeline and introduce advanced evaluation methods, including label set level metrics and a multilabel confusion matrix for error analysis. Finetuned GPT-3.5 achieved top performance with 0.94 partial match accuracy and 0.91 F1 score, while GPT-4.5 with guided prompting showed superior performance across label sets, including rare or minority label sets, indicating a more balanced and robust performance. Our findings reveal systematic error patterns, such as the conflation of SI and SA, and highlight the models tendency toward cautious over labeling. This work not only demonstrates the feasibility of using generative AI for complex clinical classification tasks but also provides a blueprint for structuring unstructured EHR data to support large scale clinical research and evidence based medicine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。