构建医疗实体识别评估基准,推动临床NLP模型标准化测试。
Named Clinical Entity Recognition Benchmark
- 基于OMOP标准统一多源临床数据中的疾病、药物等实体
- 支持编码、试验人群筛选等场景的模型性能评估
- 提供透明可比的测评平台,适合研究者和开发者参考
本技术报告提出一个用于评估语言模型在医疗领域表现的命名临床实体识别基准。该任务旨在从临床文本中提取结构化信息,服务于自动编码、临床试验队列识别和临床决策支持等应用。基准提供标准化评测平台,评估编码器与解码器架构等多种语言模型在多个医学领域的实体识别与分类能力。所用数据集为公开可获取的临床语料,涵盖疾病、症状、药物、操作及实验室检查等实体,并依据观测性医学结局合作计划(OMOP)通用数据模型进行标准化,确保跨系统的一致性与互操作性。模型性能主要以F1分数衡量,并辅以多种评估模式以全面揭示模型表现。报告还简要分析了已评测模型的趋势与局限。该基准旨在促进临床实体识别研究的透明度、可比性与创新。
原文摘要 · Abstract (English)
This technical report introduces a Named Clinical Entity Recognition Benchmark for evaluating language models in healthcare, addressing the crucial natural language processing (NLP) task of extracting structured information from clinical narratives to support applications like automated coding, clinical trial cohort identification, and clinical decision support. The leaderboard provides a standardized platform for assessing diverse language models, including encoder and decoder architectures, on their ability to identify and classify clinical entities across multiple medical domains. A curated collection of openly available clinical datasets is utilized, encompassing entities such as diseases, symptoms, medications, procedures, and laboratory measurements. Importantly, these entities are standardized according to the Observational Medical Outcomes Partnership (OMOP) Common Data Model, ensuring consistency and interoperability across different healthcare systems and datasets, and a comprehensive evaluation of model performance. Performance of models is primarily assessed using the F1-score, and it is complemented by various assessment modes to provide comprehensive insights into model performance. The report also includes a brief analysis of models evaluated to date, highlighting observed trends and limitations. By establishing this benchmarking framework, the leaderboard aims to promote transparency, facilitate comparative analyses, and drive innovation in clinical entity recognition tasks, addressing the need for robust evaluation methods in healthcare NLP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。