测试可穿戴设备中深度学习模型在真实场景下的鲁棒性
In Shift and In Variance: Assessing the Robustness of HAR Deep Learning Models against Variability
- 分离用户、设备、位置和朝向变量,分析其对模型影响
- 发现数据分布偏移越大,模型性能下降越明显
- 建议评估模型时不仅看最高准确率,还要看泛化能力
使用可穿戴惯性测量单元(IMU)进行人体活动识别(HAR)有望推动医疗健康领域发展,实现持续健康监测、疾病预测与日常行为识别。尽管深度学习(DL)HAR模型在实验室数据上表现优异,但其在真实世界中的鲁棒性尚未得到充分验证,因训练与测试数据多局限于受控环境。本研究通过分离用户、设备、佩戴位置与朝向等变量,系统评估其对DL HAR模型的影响,并利用HARVAR与REALDISP数据集进行实验。采用最大均值差异(MMD)度量数据分布偏移,发现模型性能随变量增加显著下降。结果表明,不同变量对模型影响各异,且数据分布偏移与性能下降呈反比关系。复合变量效应被进一步分析,强调了真实场景下变异性的重要影响。MMD有效量化了分布偏移,并解释了性能下降原因。结合对变异性理解与评估,将有助于开发更鲁棒的模型及优化训练策略。未来模型评估应兼顾最大F1分数与泛化能力。
原文摘要 · Abstract (English)
Human Activity Recognition (HAR) using wearable inertial measurement unit (IMU) sensors can revolutionize healthcare by enabling continual health monitoring, disease prediction, and routine recognition. Despite the high accuracy of Deep Learning (DL) HAR models, their robustness to real-world variabilities remains untested, as they have primarily been trained and tested on limited lab-confined data. In this study, we isolate subject, device, position, and orientation variability to determine their effect on DL HAR models and assess the robustness of these models in real-world conditions. We evaluated the DL HAR models using the HARVAR and REALDISP datasets, providing a comprehensive discussion on the impact of variability on data distribution shifts and changes in model performance. Our experiments measured shifts in data distribution using Maximum Mean Discrepancy (MMD) and observed DL model performance drops due to variability. We concur that studied variabilities affect DL HAR models differently, and there is an inverse relationship between data distribution shifts and model performance. The compounding effect of variability was analyzed, and the implications of variabilities in real-world scenarios were highlighted. MMD proved an effective metric for calculating data distribution shifts and explained the drop in performance due to variabilities in HARVAR and REALDISP datasets. Combining our understanding of variability with evaluating its effects will facilitate the development of more robust DL HAR models and optimal training techniques. Allowing Future models to not only be assessed based on their maximum F1 score but also on their ability to generalize effectively
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。