arXiv:2606.17113cs.LGcs.CL2026-06

选对模型比堆大模型更重要,生物医学预训练能显著提升药物警戒因果推断效果。

The Critical Role of Model Selection in Causal Inference: A Comparative Analysis of Classification Models within the InferBERT Framework for Pharmacovigilance

  • 在InferBERT框架中对比四种分类模型,发现领域预训练模型表现最优。
  • BioBERT在两个药害监测数据集上准确率最高,且与传统信号一致性最强。
  • 模型规模越大不一定越好,针对性预训练比单纯扩大参数更关键。

区分药物不良反应的因果关系与虚假关联是药物警戒的核心挑战。InferBERT框架结合了变换器模型与Do-计算,但其性能高度依赖底层分类模型。本研究评估了InferBERT中模型选择的影响,包括:简单模型是否足够、领域特定预训练是否有益、扩展至大语言模型是否提升因果检测能力,以及后处理校准的效果。在两个基准数据集(解热镇痛药诱导急性肝衰竭和曲马多相关死亡)上,采用5折交叉验证并重复20次,对比了XGBoost(基线)、ALBERT(原InferBERT)、BioBERT(生物医学变换器)和Med-LLaMA(医学大模型)四类模型。评估指标包括准确率、校准误差(ECE)及因果术语与传统信号(PRR、ROR、EBGM)的杰卡德一致性;使用配对t检验检验显著性。结果表明,BioBERT在两个数据集上均取得最高准确率,而尽管参数量大且采用参数高效微调,Med-LLaMA表现较差。领域预训练具有决定性优势。校准可降低ECE,但对准确率和因果发现影响不一。此外,BioBERT的优越性也带来了与传统信号最强的一致性。研究显示,领域特定预训练相比简单基线和大型语言模型具有明显优势。在计算药物警戒中,投入于可管理的领域感知模型,比盲目扩大模型规模更有效。

原文摘要 · Abstract (English)

Distinguishing causal adverse drug events (ADEs) from spurious correlations remains a central challenge in pharmacovigilance. The InferBERT framework integrates transformer models with Do-calculus, but its success hinges on the underlying classification model. This study evaluates the impact of model choice in InferBERT, assessing whether simpler models suffice, if domain-specific pre-training helps, whether scaling to LLMs improves causal detection, and the effect of post-hoc calibration. We performed a comparative study on two benchmarks: Analgesics-induced Acute Liver Failure (AILF) and Tramadol-related Mortalities (TRAM). Four models were evaluated-XGBoost (baseline), ALBERT (original InferBERT), BioBERT (biomedical transformer), and Med-LLaMA (medical LLM)-using 5-fold cross-validation repeated over 20 runs. We measured accuracy, Expected Calibration Error (ECE) pre- and post-isotonic regression, and Jaccard concordance of causal terms with PRR, ROR, and EBGM; significance was tested with paired t-tests. BioBERT achieved the highest accuracy on both datasets, while Med-LLaMA underperformed despite its size and parameter-efficient fine-tuning. Domain-specific pre-training was decisive. Calibration improved ECE but had mixed effects on accuracy and causal discovery. BioBERT's superiority also yielded the strongest concordance with traditional pharmacovigilance signals. These results show that domain-specific pre-training provides a clear advantage over simpler baselines and larger LLMs. Investing in manageable, domain-aware models is more effective for computational pharmacovigilance than simply scaling model size.

因果推断药物警戒领域预训练模型选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。