用预训练模型自动识别临床试验数据共享声明,提升数据可发现性。
Classifiers of Data Sharing Statements in Clinical Trial Records
- 基于领域预训练模型分析文本型数据共享声明
- 预测人工标注标签的准确率高于原分类体系
- 适合需要大规模检索可用临床数据的研究者
来自临床试验的个体参与者数据(IPD)正越来越多地被共享以供科学再利用。然而,识别可用的IPD需对大型数据库中的文本型数据共享声明(DSS)进行解读。近年来,预训练语言模型在自然语言处理方面取得进展,为基于文本输入构建高效分类器提供了可能。我们在ClinicalTrials.gov中抽取的5,000条文本DSS子集中,评估了基于领域预训练语言模型的分类器在复现原始可用性类别及人工标注标签方面的表现。结果显示,预测人工标注标签的分类器性能优于直接学习原始可用性类别的模型。这表明文本DSS内容包含原始分类体系未涵盖的有效信息,此类分类器有助于在大型试验数据库中实现可用IPD的自动化识别。
原文摘要 · Abstract (English)
Digital individual participant data (IPD) from clinical trials are increasingly distributed for potential scientific reuse. The identification of available IPD, however, requires interpretations of textual data-sharing statements (DSS) in large databases. Recent advancements in computational linguistics include pre-trained language models that promise to simplify the implementation of effective classifiers based on textual inputs. In a subset of 5,000 textual DSS from ClinicalTrials.gov, we evaluate how well classifiers based on domain-specific pre-trained language models reproduce original availability categories as well as manually annotated labels. Typical metrics indicate that classifiers that predicted manual annotations outperformed those that learned to output the original availability categories. This suggests that the textual DSS descriptions contain applicable information that the availability categories do not, and that such classifiers could thus aid the automatic identification of available IPD in large trial databases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。