提出共享稀疏的多任务学习框架,统一处理不同类型结果。
Deep Multitask Learning for Mixed-Type Outcomes with Shared Sparsity

- 通过单调变换统一不同任务的输出类型,实现跨任务信息共享。
- 在高维生物数据中准确识别共有的重要预测因子,提升预测性能。
- 适合基因表达等含连续、二值及混合结果的研究场景。
现有多任务学习方法受限于对每项任务特定损失函数的依赖,当任务间结果类型不同时,损失难以直接比较,阻碍统一目标构建与任务间信息共享。本文提出一种多任务变换框架,允许各任务响应通过未知单调变换差异呈现。针对高维生物应用中预测变量维度随样本量增长而扩大,但仅少数变量具有信息的情况,引入任务间的共享稀疏性假设。基于平滑秩准则与组Lasso惩罚,通过共享第一层的多任务深度神经网络联合估计目标函数并识别重要变量。理论分析给出了非渐近过风险界和变量选择一致性。模拟研究显示该方法在预测与变量选择上表现优于对比方法。对包含连续、二值及混合结果的基因表达数据的分析进一步验证其能提升预测效果,并识别出具有生物学意义的共享预测因子。
原文摘要 · Abstract (English)
Most existing multitask learning approaches are limited by their reliance on task-specific loss functions tailored to the scale and type of each outcome. When outcomes differ across tasks, these losses are generally not directly comparable, which makes it difficult to formulate a unified objective and may limit information sharing across tasks. We propose a multitask transformation framework in which task-specific responses may differ through unknown monotone transformations. Motivated by high-dimensional biological applications in which the predictor dimension may diverge with the sample size while only a common subset of predictors is informative, we consider shared sparsity across tasks. Under this framework, we estimate the target functions and identify important predictors by optimizing a smoothed rank-based criterion with a group-Lasso penalty, implemented through a multitask deep neural network with a shared first layer. We establish the nonasymptotic excess-risk bounds, and variable-selection consistency for the proposed estimator. Simulation studies show that the proposed method achieves competitive prediction and variable-selection performance compared with competing approaches. Analyses of gene-expression studies with continuous, binary, and mixed outcomes further illustrate that the proposed method improves prediction and identifies biologically meaningful shared predictors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。