用多阶段大模型精准提取自杀相关社会因素,提升准确率与可解释性。
A Multi-Stage Large Language Model Framework for Extracting Suicide-Related Social Determinants of Health
- 分阶段处理文本,逐步增强社会因素提取精度。
- 相比基线模型,关键上下文召回率显著提升,小模型微调效果更优。
- 中间推理过程透明,适合医疗安全与政策制定者使用。
理解导致自杀事件的社会决定因素(SDoH)对早期干预至关重要。但数据驱动方法面临长尾分布、关键应激源识别困难及模型可解释性差等挑战。本文提出一种多阶段大语言模型框架,用于从非结构化文本中提取SDoH因素。在与BioBERT、GPT-3.5-turbo及DeepSeek-R1等先进模型对比中,本框架在整体提取任务和细粒度上下文检索任务上均表现更优。此外,微调小型专用模型可在保持或超越性能的同时显著降低推理成本。多阶段设计不仅提升了提取能力,还提供中间推理步骤,增强了模型可解释性。研究还通过自动化评估与初步用户实验验证了模型解释对人工标注效率与准确性的提升作用。结果表明,该方法能更准确、透明地识别自杀相关社会因素,有助于风险人群的早期发现与预防策略优化。
原文摘要 · Abstract (English)
Background: Understanding social determinants of health (SDoH) factors contributing to suicide incidents is crucial for early intervention and prevention. However, data-driven approaches to this goal face challenges such as long-tailed factor distributions, analyzing pivotal stressors preceding suicide incidents, and limited model explainability. Methods: We present a multi-stage large language model framework to enhance SDoH factor extraction from unstructured text. Our approach was compared to other state-of-the-art language models (i.e., pre-trained BioBERT and GPT-3.5-turbo) and reasoning models (i.e., DeepSeek-R1). We also evaluated how the model's explanations help people annotate SDoH factors more quickly and accurately. The analysis included both automated comparisons and a pilot user study. Results: We show that our proposed framework demonstrated performance boosts in the overarching task of extracting SDoH factors and in the finer-grained tasks of retrieving relevant context. Additionally, we show that fine-tuning a smaller, task-specific model achieves comparable or better performance with reduced inference costs. The multi-stage design not only enhances extraction but also provides intermediate explanations, improving model explainability. Conclusions: Our approach improves both the accuracy and transparency of extracting suicide-related SDoH from unstructured texts. These advancements have the potential to support early identification of individuals at risk and inform more effective prevention strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。