arXiv:2512.21005stat.MLcs.LG2025-12

用邻居数据预测零病例地区的疫情爆发,提升小样本环境预测精度。

Learning from Neighbors with PHIBP: Predicting Infectious Disease Dynamics in Data-Sparse Environments

  • 基于邻近区域数据借用统计效能,避免零计数干扰
  • 在零病例地区仍能生成可靠预测分布,准确率显著提升
  • 适合疫情数据稀疏的地区或公共卫生决策者使用

建模稀疏计数数据在众多科学领域中面临重大统计挑战。本文聚焦传染病预测,尤其关注历史上无病例报告的地理区域的疫情爆发预测。提出泊松层次化印度餐馆过程(PHIBP)的计算框架与实验应用,该方法在微生物组和生态学研究中已成功处理稀疏计数数据。PHIBP基于绝对丰度概念,系统地从相关区域借用统计强度,克服了相对率方法对零计数的敏感性。在传染病数据上的系列实验表明,该方法为生成一致的预测分布提供了稳健基础,并有效支持α多样性与β多样性等比较分析。本章强调算法实现与实验结果,证实该统一框架在数据稀疏环境下既能实现精准疫情预测,又能提供有意义的流行病学洞见。

原文摘要 · Abstract (English)

Modeling sparse count data, which arise across numerous scientific fields, presents significant statistical challenges. This chapter addresses these challenges in the context of infectious disease prediction, with a focus on predicting outbreaks in geographic regions that have historically reported zero cases. To this end, we present the detailed computational framework and experimental application of the Poisson Hierarchical Indian Buffet Process (PHIBP), with demonstrated success in handling sparse count data in microbiome and ecological studies. The PHIBP's architecture, grounded in the concept of absolute abundance, systematically borrows statistical strength from related regions and circumvents the known sensitivities of relative-rate methods to zero counts. Through a series of experiments on infectious disease data, we show that this principled approach provides a robust foundation for generating coherent predictive distributions and for the effective use of comparative measures such as alpha and beta diversity. The chapter's emphasis on algorithmic implementation and experimental results confirms that this unified framework delivers both accurate outbreak predictions and meaningful epidemiological insights in data-sparse settings.

传染病预测稀疏数据贝叶斯建模地理建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。