用少量标注数据高效挖掘病历文本,助力疫情早期预警。
Mining Unstructured Medical Texts With Conformal Active Learning
- 结合置信度校准与主动学习,动态筛选最有价值的待标注文本。
- 仅需200条人工标注即可达到优秀分类效果,适用于复杂任务。
- 轻量模型即可达成媲美深度学习的表现,适合本地化部署。
从电子健康记录(EHR)中提取相关信息对于识别症状和自动化流行病学监测至关重要。通过挖掘EHR中大量非结构化文本,可发现疾病暴发前兆,实现更快、更精准的公共卫生响应。本文提出的框架提供了一种灵活高效的非结构化文本数据挖掘方案,显著减少对专业人员进行大规模人工标注的需求。实验表明,该框架在仅需200条人工标注的情况下,仍能在复杂分类任务中取得优异性能。此外,该方法可配合简单轻量级模型运行,其表现不仅具有竞争力,甚至在某些情况下优于资源消耗更大的深度学习模型。这一特性不仅加快了处理速度,还保护了患者隐私——数据可在本地弱硬件上处理,无需上传至外部系统。因此,本方法为实时流行病监测提供了实用、可扩展且注重隐私的解决方案,帮助医疗机构快速有效应对新兴健康威胁。
原文摘要 · Abstract (English)
The extraction of relevant data from Electronic Health Records (EHRs) is crucial to identifying symptoms and automating epidemiological surveillance processes. By harnessing the vast amount of unstructured text in EHRs, we can detect patterns that indicate the onset of disease outbreaks, enabling faster, more targeted public health responses. Our proposed framework provides a flexible and efficient solution for mining data from unstructured texts, significantly reducing the need for extensive manual labeling by specialists. Experiments show that our framework achieving strong performance with as few as 200 manually labeled texts, even for complex classification problems. Additionally, our approach can function with simple lightweight models, achieving competitive and occasionally even better results compared to more resource-intensive deep learning models. This capability not only accelerates processing times but also preserves patient privacy, as the data can be processed on weaker on-site hardware rather than being transferred to external systems. Our methodology, therefore, offers a practical, scalable, and privacy-conscious approach to real-time epidemiological monitoring, equipping health institutions to respond rapidly and effectively to emerging health threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。