用随访信息优化问卷数据预处理,实现可解释的长期症状预测。
Global Interpretability via Automated Preprocessing: A Framework Inspired by Psychiatric Questionnaires
- 基于随访数据学习非线性预处理,稳定基线问卷项
- 最终用线性矩阵预测未来严重程度,保持全局可解释性
- 适用于精神科与非精神科纵向预测,结果清晰可追溯
精神科问卷高度依赖上下文,通常仅弱预测后续症状严重度,使预后关系难以学习。尽管灵活的非线性模型可提升预测精度,但其可解释性差会削弱临床信任。在影像和组学领域,研究者常通过预处理消除访问与仪器特异性噪声,再拟合可解释的线性模型。本文采用相同策略处理问卷数据:将预处理与预测解耦。REFINE(Redundancy-Exploiting Follow-up-Informed Nonlinear Enhancement)利用随访信息学习基线测量的非线性、条目对齐预处理,并通过线性系数矩阵将这些稳定后的基线条目映射至未来严重度。该设计限制非线性仅作用于预处理阶段,而非端到端函数类,使预后关系仍可通过单一系数矩阵实现全局可解释,而非依赖事后局部归因。在人群中,该预测器恢复了贝叶斯最优条件均值,说明全局线性可解释性无需牺牲非线性预测灵活性。实验显示,REFINE 在多种精神科与非精神科纵向预测任务中优于其他可解释方法,同时保持明确的全局归因。
原文摘要 · Abstract (English)
Psychiatric questionnaires are highly context sensitive and often only weakly predict subsequent symptom severity, which makes the prognostic relationship difficult to learn. Although flexible nonlinear models can improve predictive accuracy, their limited interpretability can erode clinical trust. In fields such as imaging and omics, investigators commonly address visit- and instrument-specific artifacts by extracting stable signal through preprocessing and then fitting an interpretable linear model. We adopt the same strategy for questionnaire data by decoupling preprocessing from prediction: REFINE (Redundancy-Exploiting Follow-up-Informed Nonlinear Enhancement) uses follow-up information to learn a nonlinear, item-aligned preprocessing of baseline measurements, and then maps these stabilized baseline items to future severity through a linear coefficient matrix. This construction constrains where nonlinearity enters rather than restricting the end-to-end predictive function class, allowing the prognostic relationship to remain globally interpretable through a single coefficient matrix rather than through post hoc local attributions. In the population, the resulting predictor recovers the Bayes-optimal conditional mean, so global linear interpretability does not require sacrificing nonlinear predictive flexibility. In experiments, REFINE outperforms other interpretable approaches while preserving clear global attribution of prognostic factors across psychiatric and non-psychiatric longitudinal prediction tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。