用可解释的稀疏模型提升多模态健康数据的个性化预测效果
Subject-Adaptive Sparse Linear Models for Interpretable Personalized Health Prediction from Multimodal Lifelog Data
- 基于全局与个体效应分离的稀疏线性建模框架
- 在10名受试者450条数据上达到黑箱模型相当性能
- 适合需透明决策的临床健康预测场景
从多模态生活日志数据中更准确地预测个性化健康结果(如睡眠质量、压力水平)具有重要的临床和实践意义。然而,当前主流方法(深度神经网络与梯度提升集成)牺牲了可解释性,且未能充分处理日志数据中的显著个体差异。为此,我们提出面向个性化的稀疏线性模型(SASL),通过普通最小二乘回归结合个体特异性交互项,系统区分全局与个体层面效应。采用基于嵌套F检验的迭代后向特征剔除法构建稀疏且统计稳健的模型。考虑到健康指标常为连续过程的离散化表示,设计了回归-阈值化方法以最大化有序目标的宏平均F1分数。针对复杂预测任务,通过置信度门控选择性融合轻量级LightGBM模型输出,在不损失可解释性的前提下提升精度。在包含约450条每日观测、10名受试者的CH-2025数据集上的评估表明,混合SASL-LightGBM框架性能媲美复杂黑箱模型,但参数更少、透明度更高,为临床医生与从业者提供清晰可操作的洞察。
原文摘要 · Abstract (English)
Improved prediction of personalized health outcomes -- such as sleep quality and stress -- from multimodal lifelog data could have meaningful clinical and practical implications. However, state-of-the-art models, primarily deep neural networks and gradient-boosted ensembles, sacrifice interpretability and fail to adequately address the significant inter-individual variability inherent in lifelog data. To overcome these challenges, we propose the Subject-Adaptive Sparse Linear (SASL) framework, an interpretable modeling approach explicitly designed for personalized health prediction. SASL integrates ordinary least squares regression with subject-specific interactions, systematically distinguishing global from individual-level effects. We employ an iterative backward feature elimination method based on nested $F$-tests to construct a sparse and statistically robust model. Additionally, recognizing that health outcomes often represent discretized versions of continuous processes, we develop a regression-then-thresholding approach specifically designed to maximize macro-averaged F1 scores for ordinal targets. For intrinsically challenging predictions, SASL selectively incorporates outputs from compact LightGBM models through confidence-based gating, enhancing accuracy without compromising interpretability. Evaluations conducted on the CH-2025 dataset -- which comprises roughly 450 daily observations from ten subjects -- demonstrate that the hybrid SASL-LightGBM framework achieves predictive performance comparable to that of sophisticated black-box methods, but with significantly fewer parameters and substantially greater transparency, thus providing clear and actionable insights for clinicians and practitioners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。