1秒内分析手机使用数据,快速识别抑郁倾向。
A Fast and Minimal System to Identify Depression Using Smartphones: Explainable Machine Learning-Based Approach
- 仅用7天应用使用数据,1秒内完成采集与分析。
- 模型准确率达82.4%,关键特征来自稳定特征选择方法。
- 通过SHAP分析揭示抑郁相关行为标志,适合资源有限地区使用。
现有基于设备的抑郁症检测系统通常需长期数据,难以实现早期预警。本研究旨在开发一种快速、极简的检测系统。我们设计了一款工具,可在1秒内(平均0.31秒,标准差1.10秒)获取用户过去7天的应用使用数据。共招募100名孟加拉国学生参与实验,利用多种机器学习模型识别抑郁状态。通过稳定特征选择法结合三种主流特征选择策略,筛选关键特征。最终,轻量梯度提升机模型在仅使用1秒内采集的数据基础上,正确识别出82.4%(n=42)的抑郁学生,精确率75%,F1分数78.5%。此外,经全面探索构建的简约堆叠模型,在每轮验证中仅使用约5个由Boruta(全相关型特征选择)选出的特征,达到最高精确率77.4%(平衡准确率77.9%)。对最优模型的SHAP分析揭示了与抑郁相关的具体行为特征。该系统因快速、轻量化,有望为欠发达和发展中地区提供有效筛查支持。研究对结果的深入讨论亦有助于推动低资源消耗系统的开发,更好理解学生群体中的抑郁状况。
原文摘要 · Abstract (English)
Background: Existing robust, pervasive device-based systems developed in recent years to detect depression require data collected over a long period and may not be effective in cases where early detection is crucial. Objective: Our main objective was to develop a minimalistic system to identify depression using data retrieved in the fastest possible time. Methods: We developed a fast tool that retrieves the past 7 days' app usage data in 1 second (mean 0.31, SD 1.10 seconds). A total of 100 students from Bangladesh participated in our study, and our tool collected their app usage data. To identify depressed and nondepressed students, we developed a diverse set of ML models. We selected important features using the stable approach, along with 3 main types of feature selection (FS) approaches. Results: Leveraging only the app usage data retrieved in 1 second, our light gradient boosting machine model used the important features selected by the stable FS approach and correctly identified 82.4% (n=42) of depressed students (precision=75%, F1-score=78.5%). Moreover, after comprehensive exploration, we presented a parsimonious stacking model where around 5 features selected by the all-relevant FS approach Boruta were used in each iteration of validation and showed a maximum precision of 77.4% (balanced accuracy=77.9%). A SHAP analysis of our best models presented behavioral markers that were related to depression. Conclusions: Due to our system's fast and minimalistic nature, it may make a worthwhile contribution to identifying depression in underdeveloped and developing regions. In addition, our detailed discussion about the implication of our findings can facilitate the development of less resource-intensive systems to better understand students who are depressed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。