检验大模型泛化能力是否普适,发现不同测试集间表现无固定关联。
Do Generalisation Results Generalise?
- 多数据集动态评估微调过程中的泛化性能。
- 控制域内表现后,多数测试集间无显著相关性。
- 结果提示单一测试集评估易误导,适合关注模型鲁棒性的研究者。
大语言模型(LLM)的分布外(OOD)泛化能力对其实际部署至关重要。以往评估通常仅基于单一分布外数据集,但实际部署时的数据偏移更为多样,单一评估可能无法准确反映模型能力。本文探究了泛化结果是否具有普适性:在微调过程中持续评估模型在多个分布外测试集上的表现,并通过偏相关分析去除域内性能影响,考察各测试集间泛化表现的相关性。以OLMo2和OPT模型为对象,发现不存在统一趋势:任意两个分布外测试集间的正负相关性强烈依赖于具体模型的选择,表明泛化性能不具备跨数据集一致性。
原文摘要 · Abstract (English)
A large language model's (LLM's) out-of-distribution (OOD) generalisation ability is crucial to its deployment. Previous work assessing LLMs' generalisation performance, however, typically focuses on a single out-of-distribution dataset. This approach may fail to precisely evaluate the capabilities of the model, as the data shifts encountered once a model is deployed are much more diverse. In this work, we investigate whether OOD generalisation results generalise. More specifically, we evaluate a model's performance across multiple OOD testsets throughout a finetuning run; we then evaluate the partial correlation of performances across these testsets, regressing out in-domain performance. This allows us to assess how correlated are generalisation performances once in-domain performance is controlled for. Analysing OLMo2 and OPT, we observe no overarching trend in generalisation results: the existence of a positive or negative correlation between any two OOD testsets depends strongly on the specific choice of model analysed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。