arXiv:2607.10633cs.LGcs.AI2026-07

XAI模型看似稳定的预测结果,可能只是因变量构造导致的假象。

Auditing Construct Overlap in Explainable Machine Learning: Evidence from Burnout-Depression Prediction Across Student Cohorts

论文配图:Auditing Construct Overlap in Explainable Machine Learning: Evidence from Burnout-Depression Prediction Across Student Cohorts
图 1 · 摘自论文原文
  • 通过残差化分析,揭示了心理量表间共线性如何误导模型重要性排序。
  • 当去除共线性影响后,原本排名靠前的焦虑项重要性大幅下降,模型性能显著降低。
  • 提出可复用的残差检验方法,适用于任何涉及相关变量的可解释性研究。

将可解释机器学习(XML)应用于复合心理健康指标时,所得跨群体稳定的风险排序,常是因变量构建方式造成的伪像。本研究以洛桑大学886名医学生(主队列,2022年)为数据基础,采用ElasticNet模型,在三个时间点共2580次纵向观测及701名非医学专业学生(八个学院)中验证,三组数据使用相同测量工具。模型始终显示特质焦虑(STAI-T)和健康满意度占主导地位,五个评估集中的前两名位置肯德尔τ相关系数均为1.0,且预测性能稳定(R²:0.41–0.49)。通过两组残差化实验——将特质焦虑对抑郁分量表(CES-D,r=0.72)进行回归残差化,或将倦怠子量表对CES-D残差化——发现:前者使模型R²从0.41降至0.16,STAI-T排名从第1跌至第6;后者则导致R²骤降至0.016。预测区间平均为35.4单位(0–100量表),相当于2.4个标准差,排除了个案部署可行性。该残差化流程是本文的核心贡献:任何结合相关预测因子与目标变量的研究,均应在解释稳定性前执行此检验。

原文摘要 · Abstract (English)

Explainable machine learning (XML) pipelines applied to composite mental health outcomes can produce apparently-robust, cross-population-stable risk hierarchies that are largely artefacts of how the outcome was constructed. We demonstrate this using an ElasticNet pipeline applied to 886 medical students at the University of Lausanne (primary cohort, 2022), validated across 2,580 longitudinal observations at three time points and 701 non-medical students from eight faculties; all three datasets share identical instruments. The pipeline produces a hierarchy in which trait anxiety and health satisfaction dominate wherever the outcome is measured, with Kendall $τ= 1.0$ for the top-two positions across all five evaluation sets and consistent transfer performance ($R^2$: 0.41-0.49). Two residualization experiments, which isolate shared variance between correlated variables via regression, reveal the mechanism: when trait anxiety (STAI-T) is residualized against the co-included depression subscale (CES-D, $r = 0.72$), model $R^2$ drops from 0.41 to 0.16 and STAI-T falls from rank 1 to rank 6; when burnout subscales are residualized against CES-D, $R^2$ collapses to 0.016. Prediction intervals average 35.4 units on a 0-100 scale (2.4 outcome standard deviations), independently ruling out individual-level deployment. The residualization protocol is the paper's transferable contribution: any XAI study combining correlated predictor and outcome constructs should apply this check before interpreting apparent stability as a finding.

可解释AI心理预测共线性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。