集成学习可提升手写病历识别准确率,且不依赖训练数据量。
Evaluation of Ensemble Learning Techniques for handwritten OCR Improvement
- 融合多个机器学习模型提升OCR识别精度。
- 集成方法使识别准确率显著提高,效果稳定。
- 适合医疗历史数据数字化场景,对数据量要求低。
2021年,Lippert教授研究组的本科项目需对历史患者病历的手写内容进行数字化,采用光学字符识别(OCR)技术。由于数据将长期使用,准确性至关重要,尤其在医疗领域更为关键。集成学习通过组合多个机器学习模型,被认为可提升现有方法的性能。本文研究了集成学习与OCR结合的应用,旨在为病历数字化创造额外价值。结果表明,集成学习能有效提高OCR准确率,不同方法均表现良好,且训练数据集大小对此无显著影响。
原文摘要 · Abstract (English)
For the bachelor project 2021 of Professor Lippert's research group, handwritten entries of historical patient records needed to be digitized using Optical Character Recognition (OCR) methods. Since the data will be used in the future, a high degree of accuracy is naturally required. Especially in the medical field this has even more importance. Ensemble Learning is a method that combines several machine learning models and is claimed to be able to achieve an increased accuracy for existing methods. For this reason, Ensemble Learning in combination with OCR is investigated in this work in order to create added value for the digitization of the patient records. It was possible to discover that ensemble learning can lead to an increased accuracy for OCR, which methods were able to achieve this and that the size of the training data set did not play a role here.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。