用66种传染病数据做迁移学习,提升疾病预测准确性。
Transfer Learning using 66 Diseases for Disease Forecasting Applications

- 跨疾病迁移学习,利用66种病的数据增强预测模型。
- 84.9%的病历序列预测性能因多源数据融合而提升。
- 数据质量关键,差异过大反而降低预测效果。
疾病预测模型通常依赖单一数据源,导致在历史数据短或噪声多时表现脆弱。近期表现优异的模型表明,整合同一疾病的多个报告系统可提升性能。另有研究进一步提出使用迁移学习,通过不同疾病的训练数据来建模目标疾病。本文在此基础上大幅扩展,基于涵盖66种传染病和多个数据流的海量数据训练机器学习模型,评估多种数据流对20种疾病预测的影响。结果发现,在绝大多数(84.9%)的时间序列与模型结构中,引入其他数据流能有效提升预测表现。然而,研究也指出数据质量至关重要——若新增数据与目标数据差异过大,反而可能降低预测性能。本工作一大贡献是构建了一个公开可获取的传染病数据数据库,供流行病预测社区共享使用。
原文摘要 · Abstract (English)
Disease forecasting models typically rely on a single data stream, making models brittle when histories are short or noisy. Recent top-performing models have shown that synthesizing multiple reporting systems for the same disease improves performance. Other recent work takes this idea a step further, using transfer learning to train a forecasting model for one disease using data from a different disease. We expand upon each of these approaches greatly, training machine learning models on data that span 66 infectious diseases and several data streams. We investigate the value of incorporating different data streams for forecasting 20 different disease data streams. We find that incorporating other data streams improves forecasting in the vast majority (84.9%) of time series and model structures considered. However, our work highlights that the quality of the added data matters, where adding data extremely different from the target data stream can sometimes degrade forecast performance. A major contribution of this work is in compiling a publicly-available database of data for use by the infectious disease forecasting community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。