arXiv:2505.02257stat.MEcs.LG2025-05被引 1

用联邦学习解决死亡原因分类中的数据分布偏移问题

Bayesian Federated Cause-of-Death Classification and Quantification Under Distribution Shift

  • 基于贝叶斯联邦学习,不共享原始数据即可联合建模
  • 在无标签或少标签场景下仍能准确分类死因并估算死亡比例
  • 兼容现有算法,适合隐私敏感的公共卫生监测系统

在缺乏医学认证死因信息的地区,口头尸检(VA)通过访谈照护者来确定死因,其数据常使用概率算法分析。然而,不同人群间的数据分布差异会导致现有算法性能下降。多数现有方法依赖集中式训练,需访问全部训练数据,这在隐私和物流上常不可行。本文提出一种新的贝叶斯联邦学习(BFL)框架,避免跨机构数据共享。该方法可在目标区域仅具有限或无本地标注数据的情况下,实现个体级死因分类与群体层面特定死因死亡率的量化。所提框架模块化、计算高效,可与多种现有VA算法作为基模型集成,便于在真实死亡监测系统中灵活部署。我们在两个真实世界VA数据集上进行大量实验,验证了在不同分布偏移程度下的性能。结果表明,BFL显著优于单域基线模型,在多数情况下表现相当于或优于联合建模。

原文摘要 · Abstract (English)

In regions lacking medically certified causes of death, verbal autopsy (VA) is a widely used tool to ascertain the cause of death through interviews with caregivers. Data collected by VAs are often analyzed using probabilistic algorithms. The performance of these algorithms often degrades due to distribution shift across populations. Most existing VA algorithms rely on centralized training, requiring full access to training data for joint modeling. This can be infeasible due to privacy and logistical constraints. In this paper, we propose a novel Bayesian Federated Learning (BFL) framework that avoids data sharing across multiple training sources. Our method supports individual-level cause-of-death classification and population-level quantification of cause-specific mortality fractions in a target domain with limited or no local labeled data. The proposed framework is modular, computationally efficient, and compatible with a wide range of existing VA algorithms as base models, facilitating flexible deployment in real-world mortality surveillance systems. We validate the performance of BFL through extensive experiments on two real-world VA datasets under varying levels of distribution shift scenarios. Our results show that BFL significantly outperforms single-domain base models and performs comparably to or better than joint modeling.

联邦学习死亡原因分类医疗数据分布偏移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。