arXiv:2602.19855cs.CL2026-02

用语义嵌入融合统计信号,自动发现临床试验中的安全风险关联

SHIELD: Semantic Heterogeneity Integrated Embedding for Latent Discovery in Clinical Trial Safety Signals

  • 结合差异性分析与医学术语语义聚类,挖掘不良事件关联
  • 通过信息论指标与贝叶斯收缩计算信号强度,识别显著风险
  • 生成可解释的聚类图谱,适合药物安全评估人员使用

我们提出SHIELD,一种用于临床试验中自动化整合安全信号检测的新方法。该方法将差异性分析与不良事件(AE)术语的语义聚类相结合,基于MedDRA术语嵌入进行处理。针对每个不良事件,通过经验贝叶斯收缩推导效应量,计算信息论意义上的差异性度量(信息分量)。利用信号强度加权语义相似性构建效用矩阵,随后进行谱嵌入与聚类,以识别相关不良事件组。再借助大语言模型为聚类结果标注综合征级摘要标签,形成治疗相关安全特征的网络图与层次树结构。本方法可在单臂、双臂或多臂试验中实现安全信号检测与对比分析。在真实临床试验案例中,成功恢复已知安全信号,并生成可解释的聚类摘要。该工作将统计信号检测与现代自然语言处理结合,提升临床试验安全性评估与因果推断能力。

原文摘要 · Abstract (English)

We present SHIELD, a novel methodology for automated and integrated safety signal detection in clinical trials. SHIELD combines disproportionality analysis with semantic clustering of adverse event (AE) terms applied to MedDRA term embeddings. For each AE, the pipeline computes an information-theoretic disproportionality measure (Information Component) with effect size derived via empirical Bayesian shrinkage. A utility matrix is constructed by weighting semantic term-term similarities by signal magnitude, followed by spectral embedding and clustering to identify groups of related AEs. Resulting clusters are annotated with syndrome-level summary labels using large language models, yielding a coherent, data-driven representation of treatment-associated safety profiles in the form of a network graph and hierarchical tree. We implement the SHIELD framework in the context of a single-arm incidence summary, to compare two treatment arms or for the detection of any treatment effect in a multi-arm trial. We illustrate its ability to recover known safety signals and generate interpretable, cluster-based summaries in a real clinical trial example. This work bridges statistical signal detection with modern natural language processing to enhance safety assessment and causal interpretation in clinical trials.

临床安全语义聚类信号检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。