用代理任务预测大模型新能力,无需等模型长大就能提前知道它能不能用工具。
Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need
- 通过性能差异筛选与目标任务相关的代理任务
- 小模型集成验证后选出最优代理任务,预测结果与实际高度相关
- 适合想提前评估大模型潜力的研究者和开发者
尽管缩放定律通过在小模型或早期模型上实验来优化大语言模型(LLMs)的训练配置,但无法预测新兴能力,因为这些能力在小模型中并不存在。为此,我们提出一种利用代理任务预测新兴能力的方法。首先,基于多个模型在不同任务上的表现差异,建立目标任务与候选任务之间的相关性度量。随后,通过小模型集成对候选任务进行鲁棒性验证,筛选出最合适的代理任务。最终,通过整合这些代理任务的评估结果,推断目标任务的性能。在工具使用能力的案例研究中,预测性能与实际表现展现出强相关性,验证了方法的有效性。
原文摘要 · Abstract (English)
While scaling laws optimize training configurations for large language models (LLMs) through experiments on smaller or early-stage models, they fail to predict emergent abilities due to the absence of such capabilities in these models. To address this, we propose a method that predicts emergent abilities by leveraging proxy tasks. We begin by establishing relevance metrics between the target task and candidate tasks based on performance differences across multiple models. These candidate tasks are then validated for robustness with small model ensembles, leading to the selection of the most appropriate proxy tasks. The predicted performance on the target task is then derived by integrating the evaluation results of these proxies. In a case study on tool utilization capabilities, our method demonstrated a strong correlation between predicted and actual performance, confirming its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。