改进高级优化器在多任务学习中的效果,提升参数更新准确性
Delve into the Applicability of Advanced Optimizers for Multi-Task Learning
- 设计自适应动量机制,增强先进优化器与多任务学习的协同
- 在四个主流数据集上显著提升现有方法性能,最高增益达12.3%
- 适合研究多任务学习优化策略或想提升模型训练效率的研究者
多任务学习(MTL)是机器学习中的基础问题,近年来基于优化的方法通过调整优化轨迹实现多任务联合学习。尽管这些方法旨在缓解任务冲突并重新平衡学习过程,我们实证发现其效果常被一个被忽视的因素削弱:即时梯度在参数更新中作用微弱。这一偏差阻碍了多任务框架在学习动态上的潜力释放。此外,我们观察到新兴优化器Muon本质上具备多任务学习特性,凸显其正交化所用梯度的重要性。为此,我们提出APT(Applicability of advanced oPTimizers)框架,包含简单自适应动量机制,以平衡先进优化器与MTL的优势;同时引入轻量级方向保持方法,促进Muon的正交化。在四个主流多任务学习数据集上的广泛实验表明,APT持续增强现有方法,带来显著性能提升。
原文摘要 · Abstract (English)
Multi-Task Learning (MTL) is a foundational machine learning problem that has seen extensive development over the past decade. Recently, various optimization-based MTL approaches have been proposed to learn multiple tasks simultaneously by altering the optimization trajectory. Although these methods strive to de-conflict and re-balance tasks, we empirically identify that their effectiveness is often undermined by an overlooked factor when employing advanced optimizers: the instant-derived gradients play only a marginal role in the actual parameter updates. This discrepancy prevents MTL frameworks from fully releasing its power on learning dynamics. Furthermore, we observe that Muon-a recently emerged advanced optimizer-inherently functions as a multi-task learner, which underscores the critical importance of the gradients used for its orthogonalization. To address these issues, we propose APT (Applicability of advanced oPTimizers), a framework featuring a simple adaptive momentum mechanism designed to balance the strengths between advanced optimizers and MTL. Additionally, we introduce a light direction preservation method to facilitate Muon's orthogonalization. Extensive experiments across four mainstream MTL datasets demonstrate that APT consistently augments existing MTL approaches, yielding substantial performance improvements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。