用语义相似度提升药物安全数据聚类,效果优于传统方法
Ontology-based Semantic Similarity Measures for Clustering Medical Concepts in Drug Safety
- 基于本体信息含量的相似度算法更准确
- 新方法在聚类任务中F1达0.403,显著优于路径法
- 适合药物警戒领域研究人员快速发现风险信号
语义相似度度量(SSMs)在生物医学研究中广泛应用,但在药物警戒中仍较少使用。本研究评估了六种基于本体的SSMs在药物安全数据中对MedDRA优选术语(PTs)进行聚类的效果。利用统一医学语言系统(UMLS),我们评估各方法将PTs围绕医学意义中心点分组的能力。开发了一个高通量框架,支持Java API及Python、R接口,实现大规模相似度计算。结果表明,路径法表现中等,WUPALMER和LCH的F1分数分别为0.36和0.28;而基于内在信息含量(IC)的方法,特别是INTRINSIC-LIN和SOKAL,始终表现更优,F1分数达到0.403。经专家评审和标准MedDRA查询(SMQs)验证,结果表明基于IC的SSMs在提升药物警戒工作流方面具有潜力,可增强早期信号检测并减少人工审核负担。
原文摘要 · Abstract (English)
Semantic similarity measures (SSMs) are widely used in biomedical research but remain underutilized in pharmacovigilance. This study evaluates six ontology-based SSMs for clustering MedDRA Preferred Terms (PTs) in drug safety data. Using the Unified Medical Language System (UMLS), we assess each method's ability to group PTs around medically meaningful centroids. A high-throughput framework was developed with a Java API and Python and R interfaces support large-scale similarity computations. Results show that while path-based methods perform moderately with F1 scores of 0.36 for WUPALMER and 0.28 for LCH, intrinsic information content (IC)-based measures, especially INTRINSIC-LIN and SOKAL, consistently yield better clustering accuracy (F1 score of 0.403). Validated against expert review and standard MedDRA queries (SMQs), our findings highlight the promise of IC-based SSMs in enhancing pharmacovigilance workflows by improving early signal detection and reducing manual review.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。