针对中文病名归一化数据不足问题,提出一套增强方法提升模型性能。
Data Augmentation Techniques for Chinese Disease Name Normalization
- 基于多种技术构建数据增强流程,扩充训练样本
- 在少样本场景下显著提升多个基线模型表现
- 适合医疗文本处理与小样本学习研究者使用
病名归一化是医疗领域的重要任务,将不同格式的疾病名称映射为标准名称,是智能医疗系统中各类疾病相关功能的基础。然而,现有系统面临的最大挑战是训练数据严重不足。为此,本文提出一种新颖的数据增强方法,包含一系列数据增强技术及配套模块,以缓解数据稀缺问题。通过大量实验验证,所提方法在多种基线模型和训练目标下均表现出显著性能提升,尤其在训练数据有限的场景下效果突出。
原文摘要 · Abstract (English)
Disease name normalization is an important task in the medical domain. It classifies disease names written in various formats into standardized names, serving as a fundamental component in smart healthcare systems for various disease-related functions. Nevertheless, the most significant obstacle to existing disease name normalization systems is the severe shortage of training data. Consequently, we present a novel data augmentation approach that includes a series of data augmentation techniques and some supporting modules to help mitigate the problem. Through extensive experimentation, we illustrate that our proposed approach exhibits significant performance improvements across various baseline models and training objectives, particularly in scenarios with limited training data
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。