提出黑箱方法精准分割电子病历,准确率超98%。
Black-Box Segmentation of Electronic Medical Records
- 用简单句向量+神经网络实现黑箱分段,无需依赖特定结构
- 在多种格式数据上达到98%以上准确率,通用性强
- 适合医疗NLP初学者和需要稳定分段的系统开发者
电子病历(EMR)包含患者大部分医疗信息,是构建自动化医疗系统的重要资源。然而,当前大多数自然语言处理(NLP)研究在处理EMR时,因章节分割不准确而受到影响。同时,针对EMR精确分段的研究仍不足,导致段落结构中的信息未被充分挖掘。本文聚焦于EMR分段问题,提出一种基于简单句向量模型与神经网络的黑箱分段方法,并配合合适的训练策略。为实现通用适应性,模型在包含不同标题格式的数据集上进行训练。在多个先进深度学习方法的对比中,该方法在多种测试数据上均取得最佳分段准确率(超过98%),且在合理训练语料下表现稳定。
原文摘要 · Abstract (English)
Electronic medical records (EMRs) contain the majority of patients' healthcare details. It is an abundant resource for developing an automatic healthcare system. Most of the natural language processing (NLP) studies on EMR processing, such as concept extraction, are adversely affected by the inaccurate segmentation of EMR sections. At the same time, not enough attention has been given to the accurate sectioning of EMRs. The information that may occur in section structures is unvalued. This work focuses on the segmentation of EMRs and proposes a black-box segmentation method using a simple sentence embedding model and neural network, along with a proper training method. To achieve universal adaptivity, we train our model on the dataset with different section headings formats. We compare several advanced deep learning-based NLP methods, and our method achieves the best segmentation accuracies (above 98%) on various test data with a proper training corpus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。