开发开源工具medkit,简化临床文本表型分析流程
Facilitating phenotyping from clinical texts: the medkit library
- 用可复用的软件模块构建文本处理流水线
- 支持从电子病历中高效提取表型信息
- 适合医疗数据研究者快速搭建分析流程
表型分析旨在通过算法识别具有特定(可能复杂)特征或疾病状态的个体,通常基于电子健康记录(EHR)集合。由于大量临床信息存在于文本中,从文本中进行表型分析在依赖EHR二次利用的研究中至关重要。然而,临床文本内容与形式的高度异质性及专业性使该任务极为繁琐,成为观察性研究中时间与成本的瓶颈。为此,我们开发了一个名为medkit的开源Python库,支持通过可复用的软件组件(称为medkit操作)构建数据处理流水线。除核心功能外,我们还分享了已开发的操作和流水线,邀请表型分析社区共同使用与扩展。medkit可在https://github.com/medkit-lib/medkit获取。
原文摘要 · Abstract (English)
Phenotyping consists in applying algorithms to identify individuals associated with a specific, potentially complex, trait or condition, typically out of a collection of Electronic Health Records (EHRs). Because a lot of the clinical information of EHRs are lying in texts, phenotyping from text takes an important role in studies that rely on the secondary use of EHRs. However, the heterogeneity and highly specialized aspect of both the content and form of clinical texts makes this task particularly tedious, and is the source of time and cost constraints in observational studies. To facilitate the development, evaluation and reproductibility of phenotyping pipelines, we developed an open-source Python library named medkit. It enables composing data processing pipelines made of easy-to-reuse software bricks, named medkit operations. In addition to the core of the library, we share the operations and pipelines we already developed and invite the phenotyping community for their reuse and enrichment. medkit is available at https://github.com/medkit-lib/medkit
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。