提出双编码器框架,统一处理多任务学习中的共享与异质信息。
Multi-task Learning for Heterogeneous Data via Integrating Shared and Task-Specific Encodings
- 设计共享与任务专用双编码器,建模跨任务共性与个性特征。
- 在五种癌症的PDX数据上,预测肿瘤倍增时间表现优于现有方法。
- 适用于医疗、生物等存在分布和后验异质性的多任务场景。
多任务学习(MTL)已成为同时解决多个学习任务的重要工具,广泛应用于医疗、营销和生物医学研究等领域。为实现任务间高效信息共享,必须兼顾共享信息与异质信息。然而,现有方法难以在统一框架下处理分布异质性和后验异质性等多重异质形式。本文提出一种双编码器框架,为每个任务构建异质潜在因子空间:通过任务共享编码器捕捉跨任务共性,任务专用编码器保留各任务特异性。此外,探索潜在因子系数的内在相似结构,实现对后验异质性的自适应整合。提出一种交替优化任务专用与共享编码器及系数的统一算法。理论上,基于局部Rademacher复杂度分析了所提MTL方法的过失风险界,并应用于新但相关任务。模拟实验表明,该方法在多种设置下均优于现有数据整合方法;在五种不同癌症类型的PDX数据上,对肿瘤倍增时间的预测性能显著提升。
原文摘要 · Abstract (English)
Multi-task learning (MTL) has become an essential machine learning tool for addressing multiple learning tasks simultaneously and has been effectively applied across fields such as healthcare, marketing, and biomedical research. However, to enable efficient information sharing across tasks, it is crucial to leverage both shared and heterogeneous information. Despite extensive research on MTL, various forms of heterogeneity, including distribution and posterior heterogeneity, present significant challenges. Existing methods often fail to address these forms of heterogeneity within a unified framework. In this paper, we propose a dual-encoder framework to construct a heterogeneous latent factor space for each task, incorporating a task-shared encoder to capture common information across tasks and a task-specific encoder to preserve unique task characteristics. Additionally, we explore the intrinsic similarity structure of the coefficients corresponding to learned latent factors, allowing for adaptive integration across tasks to manage posterior heterogeneity. We introduce a unified algorithm that alternately learns the task-specific and task-shared encoders and coefficients. In theory, we investigate the excess risk bound for the proposed MTL method using local Rademacher complexity and apply it to a new but related task. Through simulation studies, we demonstrate that the proposed method outperforms existing data integration methods across various settings. Furthermore, the proposed method achieves superior predictive performance for time to tumor doubling across five distinct cancer types in PDX data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。