用真实世界数据预测动物用药风险,识别致命副作用与残留隐患。
Predictive Modeling and Explainable AI for Veterinary Safety Profiles, Residue Assessment, and Health Outcomes Using Real-World Data and Physicochemical Properties
- 融合药物理化性质与医学报告,构建百万级安全事件预测模型。
- 集成模型达95%精确率与召回率,显著提升对致死事件的识别能力。
- 通过可解释AI揭示心肺疾病和药物特性等关键风险因素,适合监管决策参考。
食品动物用药安全关乎动物福利与人类食品安全。不良事件可能预示意外药代或毒代效应,增加食物链中违规残留风险。本研究基于美国FDA开放数据库1987年至2025年第一季度的约128万份报告,构建分类框架以预测死亡与恢复结果。通过VeDDRA本体标准化不良事件,整合药物理化属性,并完成数据归一化、缺失值填补及高基数特征降维。评估了随机森林、CatBoost、XGBoost、ExcelFormer及大语言模型(Gemma 3-27B、Phi 3-12B)等监督学习方法。针对类别不平衡问题,采用欠采样与过采样策略,重点优化致死事件的召回率。集成方法(投票、堆叠)与CatBoost表现最佳,实现0.95的精确率、召回率与F1分数。引入基于平均不确定性边界(AUM)的伪标签机制,提升模型在低频类别中的检测性能,尤其改善ExcelFormer与XGBoost表现。通过SHAP分析发现,肺部、心脏及支气管疾病、动物年龄性别等特征与致死结果密切相关,具备生物学合理性。整体框架表明,结合严谨数据工程、先进机器学习与可解释人工智能,可实现高精度、可解释的兽医安全结局预测,支持FARAD使命,助力高风险药物-事件组合的早期识别,强化残留风险评估,支撑监管与临床决策。
原文摘要 · Abstract (English)
The safe use of pharmaceuticals in food-producing animals is vital to protect animal welfare and human food safety. Adverse events (AEs) may signal unexpected pharmacokinetic or toxicokinetic effects, increasing the risk of violative residues in the food chain. This study introduces a predictive framework for classifying outcomes (Death vs. Recovery) using ~1.28 million reports (1987-2025 Q1) from the U.S. FDA's OpenFDA Center for Veterinary Medicine. A preprocessing pipeline merged relational tables and standardized AEs through VeDDRA ontologies. Data were normalized, missing values imputed, and high-cardinality features reduced; physicochemical drug properties were integrated to capture chemical-residue links. We evaluated supervised models, including Random Forest, CatBoost, XGBoost, ExcelFormer, and large language models (Gemma 3-27B, Phi 3-12B). Class imbalance was addressed, such as undersampling and oversampling, with a focus on prioritizing recall for fatal outcomes. Ensemble methods(Voting, Stacking) and CatBoost performed best, achieving precision, recall, and F1-scores of 0.95. Incorporating Average Uncertainty Margin (AUM)-based pseudo-labeling of uncertain cases improved minority-class detection, particularly in ExcelFormer and XGBoost. Interpretability via SHAP identified biologically plausible predictors, including lung, heart, and bronchial disorders, animal demographics, and drug physicochemical properties. These features were strongly linked to fatal outcomes. Overall, the framework shows that combining rigorous data engineering, advanced machine learning, and explainable AI enables accurate, interpretable predictions of veterinary safety outcomes. The approach supports FARAD's mission by enabling early detection of high-risk drug-event profiles, strengthening residue risk assessment, and informing regulatory and clinical decision-making.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。