用多个弱模型的组合,让小模型学会做超人类的任务。
EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles
- 用多个弱模型在人类数据上迭代集成,提升整体能力
- 在分布内和分布外数据上分别提升4%和6%表现
- 适合想用小模型训练大模型的研究者
随着大语言模型(LLMs)接近甚至超越人类水平,亟需方法利用仅接触人类水平数据的小模型来有效监督和增强这些强大模型。我们提出一种新方法 EnsemW2S,通过在相同有限的人类数据上训练多个弱专家,并采用逐标记级集成策略,持续优化弱模型性能,从而显著提升其协同监督强学生模型的能力。我们在分布内(ID)和分布外(OOD)数据集上评估了弱专家集成与后续强学生模型的泛化性能,特别将题目难度作为分布偏移的新维度。实验结果表明,弱专家在 ID 和 OOD 数据上分别提升 4% 和 6%,学生模型在对应数据上分别提升 3.2% 和 2.28%,验证了该方法在弱到强泛化任务中的有效性。
原文摘要 · Abstract (English)
With Large Language Models (LLMs) rapidly approaching and potentially surpassing human-level performance, it has become imperative to develop approaches capable of effectively supervising and enhancing these powerful models using smaller, human-level models exposed to only human-level data. We address this critical weak-to-strong (W2S) generalization challenge by proposing a novel method aimed at improving weak experts, by training on the same limited human-level data, enabling them to generalize to complex, super-human-level tasks. Our approach, called **EnsemW2S**, employs a token-level ensemble strategy that iteratively combines multiple weak experts, systematically addressing the shortcomings identified in preceding iterations. By continuously refining these weak models, we significantly enhance their collective ability to supervise stronger student models. We extensively evaluate the generalization performance of both the ensemble of weak experts and the subsequent strong student model across in-distribution (ID) and out-of-distribution (OOD) datasets. For OOD, we specifically introduce question difficulty as an additional dimension for defining distributional shifts. Our empirical results demonstrate notable improvements, achieving 4%, and 3.2% improvements on ID datasets and, upto 6% and 2.28% on OOD datasets for experts and student models respectively, underscoring the effectiveness of our proposed method in advancing W2S generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。