跨数据集迁移学习可提升帕金森患者步冻结检测的公平性
Evidence for Phenotype-Driven Disparities in Freezing of Gait Detection and Approaches to Bias Mitigation
- 用多中心数据评估穿戴设备模型在不同表型和人群中的检测偏差
- 表型差异导致显著不公平(DPR、EOR均低于0.8)
- 传统方法无效,跨站点迁移学习显著改善公平性和准确率
步冻结(FOG)是帕金森病(PD)的严重症状,常导致受伤跌倒。近期基于可穿戴设备的人体活动识别(HAR)技术使FOG检测成为可能,但模型的偏见与公平性仍缺乏研究。偏见指系统性误差导致不平等结果,公平性则要求在不同人群间表现一致。有偏模型可能系统性忽视特定表型或人口学特征的患者,加剧医疗差距。本研究使用多中心数据集,系统评估了先进HAR模型在不同表型和人口学特征下的偏见与公平性,采用四种缓解策略:传统方法(阈值优化、对抗去偏)与迁移学习方法(多站点迁移、微调大预训练模型)。公平性通过人口均等比(DPR)和等化几率比(EOR)量化。结果显示,模型在年龄、性别、病程及关键的FOG表型上均存在显著偏见(DPR & EOR < 0.8)。表型特异性偏见尤为严重,因震颤型与僵直型需不同临床管理。传统方法效果差:阈值优化(DPR=-0.126, EOR=+0.063)、对抗去偏(DPR=-0.008, EOR=-0.001)改善有限。而多站点迁移学习显著提升公平性(DPR=+0.037, p<0.01;EOR=+0.045, p<0.01)和性能(F1-score=+0.020, p<0.05)。跨多样化数据集的迁移学习对构建公平可靠的HAR模型至关重要,确保所有帕金森患者都能受益于可穿戴监测。
原文摘要 · Abstract (English)
Freezing of gait (FOG) is a debilitating symptom of Parkinson's disease (PD) and a common cause of injurious falls. Recent advances in wearable-based human activity recognition (HAR) enable FOG detection, but bias and fairness in these models remain understudied. Bias refers to systematic errors leading to unequal outcomes, while fairness refers to consistent performance across subject groups. Biased models could systematically underserve patients with specific FOG phenotypes or demographics, potentially widening care disparities. We systematically evaluated bias and fairness of state-of-the-art HAR models for FOG detection across phenotypes and demographics using multi-site datasets. We assessed four mitigation approaches: conventional methods (threshold optimization and adversarial debiasing) and transfer learning approaches (multi-site transfer and fine-tuning large pretrained models). Fairness was quantified using demographic parity ratio (DPR) and equalized odds ratio (EOR). HAR models exhibited substantial bias (DPR & EOR < 0.8) across age, sex, disease duration, and critically, FOG phenotype. Phenotype-specific bias is particularly concerning as tremulous and akinetic FOG require different clinical management. Conventional bias mitigation methods failed: threshold optimization (DPR=-0.126, EOR=+0.063) and adversarial debiasing (DPR=-0.008, EOR=-0.001) showed minimal improvement. In contrast, transfer learning from multi-site datasets significantly improved fairness (DPR=+0.037, p<0.01; EOR=+0.045, p<0.01) and performance (F1-score=+0.020, p<0.05). Transfer learning across diverse datasets is essential for developing equitable HAR models that reliably detect FOG across all patient phenotypes, ensuring wearable-based monitoring benefits all individuals with PD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。