用多任务学习让大模型自动标出自杀风险文本证据,提升可解释性。
Evidence-Driven Marker Extraction for Social Media Suicide Risk Detection
- 用Mistral-7B模型同时识别风险标记和分类等级,联合训练提升效率。
- 在CLPsych数据集上,标记定位准确率优于微调大模型和传统方法。
- 适合临床辅助系统开发,帮助医生快速定位高风险用户文本证据。
从社交媒体文本中早期发现自杀风险对及时干预至关重要。尽管大语言模型(LLMs)在此领域展现出潜力,但在可解释性和计算效率方面仍存在挑战。本文提出一种名为证据驱动大模型(ED-LLM)的新方法,用于临床标记提取与自杀风险分类。ED-LLM采用多任务学习框架,基于Mistral-7B模型联合训练,以识别临床标记片段并分类自杀风险等级。该证据驱动策略通过明确标注支持风险判断的文本证据,增强结果可解释性。在CLPsych数据集上的评估表明,ED-LLM在风险分类任务中表现具有竞争力,在临床标记片段识别方面显著优于基线方法,包括微调的大模型、传统机器学习及提示工程方法。结果验证了多任务学习在可解释且高效的基于大模型的自杀风险评估中的有效性,为临床相关应用铺平道路。
原文摘要 · Abstract (English)
Early detection of suicide risk from social media text is crucial for timely intervention. While Large Language Models (LLMs) offer promising capabilities in this domain, challenges remain in terms of interpretability and computational efficiency. This paper introduces Evidence-Driven LLM (ED-LLM), a novel approach for clinical marker extraction and suicide risk classification. ED-LLM employs a multi-task learning framework, jointly training a Mistral-7B based model to identify clinical marker spans and classify suicide risk levels. This evidence-driven strategy enhances interpretability by explicitly highlighting textual evidence supporting risk assessments. Evaluated on the CLPsych datasets, ED-LLM demonstrates competitive performance in risk classification and superior capability in clinical marker span identification compared to baselines including fine-tuned LLMs, traditional machine learning, and prompt-based methods. The results highlight the effectiveness of multi-task learning for interpretable and efficient LLM-based suicide risk assessment, paving the way for clinically relevant applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。