用数据量变化的多任务学习曲线,更本质地刻画迁移效果。
Characterization of Transfer Using Multi-task Learning Curves
- 通过增加样本量而非调整模型,分析迁移效应
- 多任务学习曲线能更好捕捉迁移效果,揭示成对与上下文迁移
- 适合研究基础模型迁移机制的科研人员
迁移效应既体现在固定数据集上的训练过程,也体现在随数据累积的归纳推理中。我们假设,通过增加样本量来扰动数据集,比通过梯度更新扰动模型,能提供一种互补且更根本的迁移效应表征。为此,我们定量建模迁移效应,使用多任务学习曲线模拟不同样本规模下的归纳性能。提出一种高效方法近似多任务学习曲线,类似于训练阶段的任务亲和分组方法。对比统计与计算方法,结果表明前者计算成本更高但统计功效更强、适用范围更广。在基准药物-靶点相互作用数据集上评估,结果显示学习曲线能更准确捕捉多任务学习效应,其扩展形式可区分基础模型中的成对与上下文迁移效应。
原文摘要 · Abstract (English)
Transfer effects manifest themselves both during training using a fixed data set and in inductive inference using accumulating data. We hypothesize that perturbing the data set by including more samples, instead of perturbing the model by gradient updates, provides a complementary and more fundamental characterization of transfer effects. To capture this phenomenon, we quantitatively model transfer effects using multi-task learning curves approximating the inductive performance over varying sample sizes. We describe an efficient method to approximate multi-task learning curves analogous to the Task Affinity Grouping method applied during training. We compare the statistical and computational approaches to transfer, which indicates considerably higher compute costs for the previous but better power and broader applicability. Evaluations are performed using a benchmark drug-target interaction data set. Our results show that learning curves can better capture the effects of multi-task learning and their multi-task extensions can delineate pairwise and contextual transfer effects in foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。