arXiv:2607.16554cs.LGcs.AI2026-07中稿 · 42nd Conference on…

揭示多任务学习中容量与冗余的权衡机制,提升共享性能

Capacity and Redundancy Trade-offs in Multi-Task Learning

论文配图:Capacity and Redundancy Trade-offs in Multi-Task Learning
图 1 · 摘自论文原文
  • 构建容量-冗余框架,分解任务预测信息以量化干扰
  • 发现聚类共享优于全局共享的充要条件,验证梯度相似性可表征任务冗余
  • 实证表明聚类LoRA显著降低干扰项,效果优于随机划分

多任务学习中负迁移常被视为优化副作用,但也可归因于共享容量有限和任务冗余不足。本文提出容量-冗余(CR)恒等式,将各任务预测信息之和分解为包含标签冗余(通过总相关性TC衡量)的联合预测信息,以及由共享表示无法消除的残差耦合项Δ。进一步证明:(i) 聚类间隙分解给出聚类共享优于全局共享的充要条件;(ii) 在高斯多任务模型中建立梯度-TC桥梁,正式支持梯度余弦相似性作为冗余排序代理。实验上,通过验证集残差相关性估计Δ,发现聚类LoRA显著降低ˆΔ,优于同规模随机划分,并在多种子置信区间下实现统计显著增益。

原文摘要 · Abstract (English)

In multi-task learning (MTL) negative transfer is often considered as an optimization artifact, but it can also be viewed as a consequence of limited shared capacity and weak task redundancy. We investigate this effect through a Capacity--Redundancy (CR) identity that decomposes the sum of per-task predictive informations into joint predictive information that includes label redundancy defined via total correlation (TC), and a residual coupling term that quantifies interference left unresolved by the shared representation. Additionally, we show two key results: (i) a clustering-gap decomposition that gives a necessary and sufficient condition for clustered sharing to outperform global sharing, and (ii) a gradient--TC bridge in a Gaussian multi-task model that formally justifies gradient cosine similarity as a proxy for redundancy ordering. Empirically, we estimate the residual coupling $Δ$ from validation residual correlations, showing that clustered LoRA substantially reduces $\widehatΔ$, outperforms size-matched random partitions, and results in statistically significant gains with multi-seed confidence intervals.

多任务学习容量分析冗余度梯度相似性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。