发现少样本翻译任务中主动学习失效,因核心假设不成立。
Testing the Assumptions of Active Learning for Translation Tasks with Few Samples
- 检验主动学习在小样本下的假设有效性
- 发现模型性能与数据信息量/多样性无关
- 适合研究少样本学习与主动学习的交叉方向
主动学习(AL)通过选择最具信息量和多样性的未标注样本进行标注,以提升模型在测试集上的表现,适用于标注成本高的场景。然而,近期研究表明,在仅使用100-500个样本的语言生成任务中,主动学习策略的表现无法超越随机采样。为理解其在极小样本下表现不佳的原因,本文系统检验了主动学习的核心假设是否成立。结果表明,主动学习所优化的信息量与多样性,均与测试集性能无显著相关性。相反,训练样本的顺序以及与预训练数据的交互作用对模型性能影响更大。这提示未来主动学习方法需重新考虑这些因素,才能在极低标注预算下有效工作。
原文摘要 · Abstract (English)
Active learning (AL) is a training paradigm for selecting unlabeled samples for annotation to improve model performance on a test set, which is useful when only a limited number of samples can be annotated. These algorithms often work by optimizing for the informativeness and diversity of the training data to be annotated. Recent work found that AL strategies fail to outperform random sampling on various language generation tasks when using 100-500 samples. To understand AL's poor performance when only using few samples, we investigate whether the core assumptions underlying AL strategies hold. We find that neither the informativeness nor diversity of the training data, which AL strategies optimize for, are correlated with test set performance. Instead, factors like the ordering of the training samples and interactions with pre-training data have a larger impact on performance. This suggests that future AL methods must take these factors into account in order to work with very few samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。