无需训练即可预测NLP模型学习曲线,降低实验成本。
Zero-Shot Performance Prediction for Probabilistic Scaling Laws
- 构建双层任务层次结构,用高斯过程建模任务间关联。
- 在3个数据集上实现零样本学习曲线预测,误差小。
- 适合资源有限但需快速评估模型性能的研究者。
预测自然语言处理(NLP)模型的学习曲线,有助于在满足特定性能目标的同时减少计算开销,并降低数据集获取与整理的成本。本文将预测任务建模为多任务学习问题,每个任务的数据按双层层级组织。通过潜在变量多输出高斯过程建模跨任务与层级的共享信息和依赖关系,支持零样本学习曲线预测,并能捕捉任务相关性。该方法可低成本构建概率性缩放定律。结合主动学习策略,可查询学习曲线以降低预测不确定性,使预测结果逼近真实缩放规律。我们在三个小规模NLP数据集上验证框架有效性,共生成最多30条学习曲线,涵盖nanoGPT、mBART和Transformer模型的双语翻译,以及不同规模M2M100模型的多语言翻译。
原文摘要 · Abstract (English)
The prediction of learning curves for Natural Language Processing (NLP) models enables informed decision-making to meet specific performance objectives, while reducing computational overhead and lowering the costs associated with dataset acquisition and curation. In this work, we formulate the prediction task as a multitask learning problem, where each task's data is modelled as being organized within a two-layer hierarchy. To model the shared information and dependencies across tasks and hierarchical levels, we employ latent variable multi-output Gaussian Processes, enabling to account for task correlations and supporting zero-shot prediction of learning curves (LCs). We demonstrate that this approach facilitates the development of probabilistic scaling laws at lower costs. Applying an active learning strategy, LCs can be queried to reduce predictive uncertainty and provide predictions close to ground truth scaling laws. We validate our framework on three small-scale NLP datasets with up to $30$ LCs. These are obtained from nanoGPT models, from bilingual translation using mBART and Transformer models, and from multilingual translation using M2M100 models of varying sizes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。