arXiv:2410.03140cs.LGcs.CL2024-10被引 2

提出新方法让模型在伪相关干扰下仍能有效利用上下文进行分类。

In-context Learning in Presence of Spurious Correlations

  • 设计新训练策略,使模型在存在伪特征时仍能从上下文学习。
  • 在多个任务上性能媲美甚至超越ERM和GroupDRO。
  • 通过多样化合成数据训练,实现对未见任务的泛化能力。

大型语言模型具备出色的上下文学习能力,即通过少量示例即可完成任务学习。近期研究已证明变压器模型可在上下文学习中完成简单回归任务。本文探索了在存在伪特征的情况下,训练用于分类任务的上下文学习者可能性。发现传统训练方式易受伪特征影响;当元训练数据仅包含单一任务时,模型会陷入任务记忆,无法利用上下文进行预测。基于此,我们提出一种新训练技术,用于特定分类任务的上下文学习者。该模型表现优异,性能可匹敌甚至超越强基线如ERM和GroupDRO,但不具跨任务泛化能力。进一步实验表明,通过在多样化合成上下文学习实例上训练,可实现对未见任务的有效泛化。

原文摘要 · Abstract (English)

Large language models exhibit a remarkable capacity for in-context learning, where they learn to solve tasks given a few examples. Recent work has shown that transformers can be trained to perform simple regression tasks in-context. This work explores the possibility of training an in-context learner for classification tasks involving spurious features. We find that the conventional approach of training in-context learners is susceptible to spurious features. Moreover, when the meta-training dataset includes instances of only one task, the conventional approach leads to task memorization and fails to produce a model that leverages context for predictions. Based on these observations, we propose a novel technique to train such a learner for a given classification task. Remarkably, this in-context learner matches and sometimes outperforms strong methods like ERM and GroupDRO. However, unlike these algorithms, it does not generalize well to other tasks. We show that it is possible to obtain an in-context learner that generalizes to unseen tasks by training on a diverse dataset of synthetic in-context learning instances.

上下文学习伪相关分类任务模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。