用预测模型提升药物疫苗不良反应报告去重准确率
A Scalable Predictive Modelling Approach to Identifying Duplicate Adverse Event Reports for Drugs and Vaccines
- 将统计匹配升级为预测建模,融合国家报告率与文本日期提取
- 疫苗去重精确率达92%,药品达54%,召回率超80%且跨国更稳定
- 适合药监机构、流行病学研究者用于大规模不良反应数据清洗
目标:提升大规模药物和疫苗不良事件报告中重复项检测的性能,实现跨国报告的一致性。背景:未关联的重复报告会干扰统计分析并误导临床判断。药物警戒依赖大型不良事件数据库发现潜在因果关系,需计算方法在大规模下识别重复项。当前最佳方法为统计记录链接,其中vigiMatch已用于世卫组织全球不良事件数据库VigiBase,是首个大规模部署的统计去重方法。但其在疫苗应用上因各国表现不一而受限。方法:本文将vigiMatch从概率记录链接扩展为预测建模,通过国家特异性报告率优化药品、疫苗及不良事件特征,从自由文本中提取时间信息,并分别为药品和疫苗训练支持向量机分类器。使用5个独立标注测试集评估召回率,随机抽样标注重复对评估精确率。结果:新方法对疫苗的精确率为92%,药品为54%,对比方法分别为41%;召回率在疫苗测试集中为80-85%,药品为40-86%,对比方法为24-53%。结论:预测建模、自由文本利用与国家特异性特征显著提升药物警戒中重复检测的性能。
原文摘要 · Abstract (English)
Objectives: To advance state-of-the-art for duplicate detection in large-scale pharmacovigilance databases and achieve more consistent performance across adverse event reports from different countries. Background: Unlinked adverse event reports referring to the same case impede statistical analysis and may mislead clinical assessment. Pharmacovigilance relies on large databases of adverse event reports to discover potential new causal associations, and computational methods are required to identify duplicates at scale. Current state-of-the-art is statistical record linkage which outperforms rule-based approaches. In particular, vigiMatch is in routine use for VigiBase, the WHO global database of adverse event reports, and represents the first statistical duplicate detection approach in pharmacovigilance deployed at scale. Originally developed for both medicines and vaccines, its application to vaccines has been limited due to inconsistent performance across countries. Methods: This paper extends vigiMatch from probabilistic record linkage to predictive modelling, refining features for medicines, vaccines, and adverse events using country-specific reporting rates, extracting dates from free text, and training separate support vector machine classifiers for medicines and vaccines. Recall was evaluated using 5 independent labelled test sets. Precision was assessed by annotating random selections of report pairs classified as duplicates. Results: Precision for the new method was 92% for vaccines and 54% for medicines, compared with 41% for the comparator method. Recall ranged from 80-85% across test sets for vaccines and from 40-86% for medicines, compared with 24-53% for the comparator method. Conclusion: Predictive modeling, use of free text, and country-specific features advance state-of-the-art for duplicate detection in pharmacovigilance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。