arXiv:2505.21959cs.LGcs.CL2025-05ACL被引 10

用多个弱模型集成提升强模型泛化能力,解决小数据难监督大模型的问题。

EnsemW2S: Enhancing Weak-to-Strong Generalization with Large Language Model Ensembles

  • 通过迭代融合多个弱模型的输出,提升整体推理能力。
  • 在分布内和分布外数据上分别取得4%和6%的性能提升。
  • 适合需要高效利用有限人类数据训练强模型的研究场景。

随着大型语言模型(LLMs)快速接近甚至超越人类水平表现,迫切需要发展能有效利用仅包含人类水平数据的小型模型来监督和增强这些强大模型的方法。本文针对这一关键的弱到强(W2S)泛化挑战,提出一种新方法——EnsemW2S,通过在相同有限的人类数据上训练多个弱专家,并采用逐标记级集成策略,持续迭代优化,系统性弥补前序迭代的缺陷。该方法显著提升了弱模型集体监督强学生模型的能力。我们在分布内(ID)与分布外(OOD)数据集上全面评估了弱专家集成及后续强学生模型的泛化性能。针对OOD,特别引入问题难度作为分布偏移的新维度。实验证明,该方法在ID数据集上使专家与学生模型分别提升4%和3.2%,在OOD数据集上分别提升最高达6%和2.28%,充分验证了其在推进W2S泛化方面的有效性。

原文摘要 · Abstract (English)

With Large Language Models (LLMs) rapidly approaching and potentially surpassing human-level performance, it has become imperative to develop approaches capable of effectively supervising and enhancing these powerful models using smaller, human-level models exposed to only human-level data. We address this critical weak-to-strong (W2S) generalization challenge by proposing a novel method aimed at improving weak experts, by training on the same limited human-level data, enabling them to generalize to complex, super-human-level tasks. Our approach, called \textbf{EnsemW2S}, employs a token-level ensemble strategy that iteratively combines multiple weak experts, systematically addressing the shortcomings identified in preceding iterations. By continuously refining these weak models, we significantly enhance their collective ability to supervise stronger student models. We extensively evaluate the generalization performance of both the ensemble of weak experts and the subsequent strong student model across in-distribution (ID) and out-of-distribution (OOD) datasets. For OOD, we specifically introduce question difficulty as an additional dimension for defining distributional shifts. Our empirical results demonstrate notable improvements, achieving 4\%, and 3.2\% improvements on ID datasets and, upto 6\% and 2.28\% on OOD datasets for experts and student models respectively, underscoring the effectiveness of our proposed method in advancing W2S generalization.

大模型弱监督模型集成泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。