arXiv:2606.02545cs.CL2026-06

用大模型增强机器学习,从急诊分诊记录中精准识别自伤行为

Transferable Self-Harm Surveillance from Emergency Department Triage Notes Using an Evidence-Augmented Machine Learning Approach

论文配图:Transferable Self-Harm Surveillance from Emergency Department Triage Notes Using an Evidence-Augmented Machine Learning Approach
图 1 · 摘自论文原文
  • 融合大模型筛选与证据提取,提升自伤检测能力
  • 跨三家医院验证,外推性能AUPRC达0.816以上
  • 可准确识别自伤方式,支持精细化监控

自伤是重大公共卫生问题,但现有依赖医院就诊记录的监测手段因诊断编码敏感性低而效果不佳。急诊科(ED)分诊记录在初次接触时即被记录,能简洁概括就诊情况,是识别自伤的潜在数据源。本文提出一种三阶段方法,通过大语言模型辅助筛选与证据抽取,增强传统机器学习在急诊分诊记录中识别自伤的能力。我们在三家澳大利亚医院评估了模型的跨机构迁移能力。内部与外部验证中,模型的AUPRC分别为0.887 ± 0.016和0.884 ± 0.012。前瞻性测试中,开发机构AUPRC为0.881 ± 0.008,两个外部机构分别达到0.879 ± 0.012和0.816 ± 0.015,且无需本地再训练。该方法还能以95%准确率识别主要自伤方式,支持超越二分类的细粒度监测。

原文摘要 · Abstract (English)

Self-harm is a major public health concern, but current surveillance relying on hospital presentations is inadequate due to the low sensitivity of diagnostic codes. Emergency Department (ED) triage notes, recorded at the initial point of contact, provide a succinct summary of presentations and an opportunity to identify self-harm. We developed a three-stage approach, augmenting traditional machine learning with large language model-based screening and evidence extraction to detect self-harm in ED triage notes. We assessed model transferability across three Australian hospitals. Our approach showed AUPRCs of 0.887 +/- 0.016 and 0.884 +/- 0.012 during internal and external validation. Prospectively, it achieved AUPRC of 0.881 +/- 0.008 at the development site, and 0.879 +/- 0.012 and 0.816 +/- 0.015 at two external sites without site-specific retraining. A key advantage of the approach is that it enables identification of the primary self-harm method with an accuracy of 95%, supporting more granular surveillance beyond binary classification.

自伤监测大模型应用医疗文本挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。