不同医院的自伤记录用词差异大,导致模型跨机构效果差。
Why Do Self-Harm Prediction Models Struggle to Generalise? Lexical and Semantic Variations in Emergency Department Triage Notes

- 对比两家医院的急诊记录,发现自伤描述用词和关键特征不一致。
- 尽管核心主题如自伤、服药相似,但关键词重要性差异显著。
- 研究提示需关注临床文本差异,提升模型跨机构泛化能力。
急诊科(ED)的自伤就诊与更高的自杀风险密切相关。自然语言处理(NLP)模型在单一医院的急诊分诊记录中检测自伤表现良好,但在跨机构应用时性能往往下降。为探究原因,我们比较了两家医院的急诊分诊记录,分析其词汇特征、高关联预测特征及显著话题。结果显示,尽管存在自伤、自中毒等一致的核心主题,但各机构在词汇表达方式和自伤相关特征的重要性上存在显著差异。这些文档书写差异与跨机构模型性能下降直接相关。研究揭示了机构间文本差异如何影响临床文本中自伤识别,并提出了提升模型泛化性的潜在方法。
原文摘要 · Abstract (English)
Self-harm presentations to emergency departments (EDs) are strongly associated with higher suicide risk. NLP models have shown robust performance in detecting self-harm from triage notes within single hospitals, yet performance often declines across institutions. To examine potential causes, we compare ED triage notes from two hospitals by analyzing lexical characteristics, highly associated predictive features, and salient topics. Our results reveal variation in lexical expression and feature importance related to self-harm across hospitals, despite consistent core themes such as self-poisoning and self-injury. These documentation differences are associated with reduced cross-site performance. Our findings provide insight into how institutional variation affects the identification of self-harm in clinical text and highlight potential methods to improve model generalisability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。