用大模型生成真实病历数据,提升租房被驱逐的健康风险识别准确率。
SynthEHR-Eviction: Enhancing Eviction SDoH Detection with LLM-Augmented Synthetic EHR Data
- 用LLM+人工标注+自动提示优化,从病历文本中提取租户驱逐信息。
- 在验证数据上达88.8%准确率,比GPT-4o还高,且节省超80%标注工作量。
- 适合医疗数据挖掘、公共卫生研究者,尤其关注社会健康因素的团队。
驱逐是重要的社会健康决定因素(SDoH),与住房不稳、失业和心理健康相关,但常出现在非结构化电子健康记录(EHR)中,极少被编码于结构化字段,限制了后续应用。我们提出SynthEHR-Eviction,一个结合大语言模型(LLM)、人机协同标注与自动提示优化(APO)的可扩展流水线,用于从临床笔记中提取驱逐状态。利用该流程,我们构建了目前最大公开的驱逐相关SDoH数据集,包含14个细粒度类别。在人工验证数据上,微调后的LLM(如Qwen2.5、LLaMA3)对驱逐的宏平均F1达88.8%,对其他SDoH为90.3%,优于GPT-4o-APO(87.8%、87.3%)、GPT-4o-mini-APO(69.1%、78.1%)和BioBERT(60.7%、68.3%)。该方法降低标注成本超80%,加速数据集构建,支持大规模驱逐检测,并可推广至其他信息抽取任务。
原文摘要 · Abstract (English)
Eviction is a significant yet understudied social determinants of health (SDoH), linked to housing instability, unemployment, and mental health. While eviction appears in unstructured electronic health records (EHRs), it is rarely coded in structured fields, limiting downstream applications. We introduce SynthEHR-Eviction, a scalable pipeline combining LLMs, human-in-the-loop annotation, and automated prompt optimization (APO) to extract eviction statuses from clinical notes. Using this pipeline, we created the largest public eviction-related SDoH dataset to date, comprising 14 fine-grained categories. Fine-tuned LLMs (e.g., Qwen2.5, LLaMA3) trained on SynthEHR-Eviction achieved Macro-F1 scores of 88.8% (eviction) and 90.3% (other SDoH) on human validated data, outperforming GPT-4o-APO (87.8%, 87.3%), GPT-4o-mini-APO (69.1%, 78.1%), and BioBERT (60.7%, 68.3%), while enabling cost-effective deployment across various model sizes. The pipeline reduces annotation effort by over 80%, accelerates dataset creation, enables scalable eviction detection, and generalizes to other information extraction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。