arXiv:2504.13116cs.LGstat.ME2025-04

用机器学习预测爱尔兰牛群疫病复发,精准识别高风险牛群。

Predicting BVD Re-emergence in Irish Cattle From Highly Imbalanced Herd-Level Data Using Machine Learning Algorithms

  • 采用随机森林与XGBoost模型处理极不平衡的牛群数据
  • 在2023年真实数据中识别出219个阳性牛群,准确率超87%
  • 可将检测范围缩小一半,适合疫病净化后期的精准监测

爱尔兰对牛病毒性腹泻(BVD)的根除计划成效显著,牧场层面的患病率从2013年的11.3%降至2023年的0.2%。随着迈向无疫状态,建立预测模型以支持靶向监测变得至关重要。本研究评估了多种机器学习算法在高度不平衡的牧场级数据上的表现,涵盖二分类与异常检测方法。通过大规模模拟实验,考察不同样本量与类别不平衡比下的模型性能,结合重采样、类别权重及敏感性、阳性预测值、F1分数和AUC等指标进行评估。随机森林与XGBoost表现最优,其中随机森林在各类场景中均实现最高敏感性和AUC;在2023年实际数据预测中,正确识别出250个阳性牛群中的219个,相比全面检测策略,可减少一半检测范围。

原文摘要 · Abstract (English)

Bovine Viral Diarrhoea (BVD) has been the focus of a successful eradication programme in Ireland, with the herd-level prevalence declining from 11.3% in 2013 to just 0.2% in 2023. As the country moves toward BVD freedom, the development of predictive models for targeted surveillance becomes increasingly important to mitigate the risk of disease re-emergence. In this study, we evaluate the performance of a range of machine learning algorithms, including binary classification and anomaly detection techniques, for predicting BVD-positive herds using highly imbalanced herd-level data. We conduct an extensive simulation study to assess model performance across varying sample sizes and class imbalance ratios, incorporating resampling, class weighting, and appropriate evaluation metrics (sensitivity, positive predictive value, F1-score and AUC values). Random forests and XGBoost models consistently outperformed other methods, with the random forest model achieving the highest sensitivity and AUC across scenarios, including real-world prediction of 2023 herd status, correctly identifying 219 of 250 positive herds while halving the number of herds that require compared to a blanket-testing strategy.

机器学习疫病预测数据不平衡畜牧业

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。