通过误差聚类与局部集成,提升多任务学习鲁棒性
Robust multi-task boosting using clustering and local ensembling
- 基于跨任务误差自动聚类,避免无关任务共享信息
- 在真实与合成数据上均显著优于单任务和传统多任务方法
- 适合存在噪声或无关任务的多任务场景,提升模型泛化能力
多任务学习(MTL)旨在通过共享相关任务的信息来提升预测性能,但传统方法在强制无关或含噪任务共享表示时常出现负向迁移。本文提出鲁棒多任务提升框架RMB-CLE,结合基于误差的任务聚类与局部集成。不同于以往固定聚类或人工设计相似性度量的方法,RMB-CLE直接从跨任务误差中推导任务间相似性,其误差可分解为函数不匹配与不可约噪声,具备理论基础以防止负向迁移。任务通过凝聚聚类自适应分组,每组内采用局部集成实现稳健的知识共享并保留任务特异性模式。实验表明,RMB-CLE在合成数据中能恢复真实聚类,并在多种真实与合成基准上持续优于多任务、单任务及基于池化的集成方法。结果证明,RMB-CLE不仅是聚类与提升的简单组合,更是一种通用且可扩展的鲁棒多任务学习新范式。
原文摘要 · Abstract (English)
Multi-Task Learning (MTL) aims to boost predictive performance by sharing information across related tasks, yet conventional methods often suffer from negative transfer when unrelated or noisy tasks are forced to share representations. We propose Robust Multi-Task Boosting using Clustering and Local Ensembling (RMB-CLE), a principled MTL framework that integrates error-based task clustering with local ensembling. Unlike prior work that assumes fixed clusters or hand-crafted similarity metrics, RMB-CLE derives inter-task similarity directly from cross-task errors, which admit a risk decomposition into functional mismatch and irreducible noise, providing a theoretically grounded mechanism to prevent negative transfer. Tasks are grouped adaptively via agglomerative clustering, and within each cluster, a local ensemble enables robust knowledge sharing while preserving task-specific patterns. Experiments show that RMB-CLE recovers ground-truth clusters in synthetic data and consistently outperforms multi-task, single-task, and pooling-based ensemble methods across diverse real-world and synthetic benchmarks. These results demonstrate that RMB-CLE is not merely a combination of clustering and boosting but a general and scalable framework that establishes a new basis for robust multi-task learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。