提出新方法提升抑郁学生缓解预测准确率,减少误判风险。
SMOTE-VAR: An Uncertainty-Aware Oversampling Method for Predicting Depression Remission in University Students
- 用高斯过程方差评估合成样本不确定性,避免生成无效数据
- 在大学生抑郁数据集上表现优于传统SMOTE,提升非缓解者识别率
- 适合临床辅助决策,帮助医生及时调整治疗方案
大学生抑郁症等常见心理问题发生率高,影响学习与生活。尽管正念、运动等干预可缓解症状,但许多患者无法实现症状缓解。若能提前识别预后不佳者,可实现更早、更精准的干预。机器学习用于预测抑郁缓解已受关注,但常面临类别不平衡问题——缓解者与未缓解者比例失衡,影响模型准确性和预测公平性。现有研究多采用流行的SMOTE过采样策略,但其可能生成不合理的少数类样本,在临床中易导致误判,延误需升级治疗的患者。本文提出一种新型过采样方法SMOTE-VAR,利用高斯过程的方差函数估计合成样本的不确定性,从而降低假阳性。在大学生抑郁数据集上的验证表明,该方法显著优于现有过采样技术,在预测缓解状态方面表现更优。通过更可靠地识别非缓解者,该方法为临床提供了一个高效计算工具,支持医生快速转向联合疗法,实现心理干预路径的个性化与优化。
原文摘要 · Abstract (English)
University students experience disproportionately high rates of common mental health conditions, such as depression, which can impair learning, social functioning, and overall well-being. Although lifestyle interventions such as mindfulness and physical activity can reduce the symptoms, many do not achieve symptomatic remission. Developing new approaches to identify students with poor outcomes could enable earlier and more targeted intervention. Machine learning (ML) methods have increasingly been used to predict remission in depressive patients. However, these ML models often suffer from class imbalance, where there may be an unequal proportion of people in the remitted group relative to the non-remitted group. This imbalance can reduce model accuracy and bias predictions. To address this, studies commonly employ the popular oversampling strategy SMOTE. However, SMOTE has a notable limitation: it may generate invalid synthetic minority samples. In a clinical context, these false positives can lead to incorrect risk stratification, potentially delaying necessary escalated care for patients unlikely to remit. In this paper, we introduce a novel and effective oversampling method that addresses this shortcoming. Our approach leverages the variance function of a Gaussian process to estimate the uncertainty of generated minority samples to reduce false positives. We validate our method on a depression dataset collected from university students and demonstrate that it is better than existing oversampling approaches in predicting remission (i.e., treatment outcome). By improving the reliable identification of non-responders, our method provides a robust computational tool to help clinicians rapidly pivot to adjunctive therapies, thereby personalizing and optimizing mental health care pathways.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。