用目标掩码预训练提升伪标签半监督文本分类效率
The Efficiency of Pre-training with Objective Masking in Pseudo Labeling for Semi-Supervised Text Classification
- 引入目标掩码预训练,增强教师模型初始性能
- 在英、瑞双语数据集上显著提升分类准确率
- 适合小样本标注场景下的高效文本分类任务
我们扩展并深入研究了Hatefi等人提出的半监督文本分类模型,该模型适用于仅少量标注文档的情况下进行分类,多数训练样本无标签。模型采用元伪标签的师生架构,由‘教师’为原始无标签数据生成伪标签以训练‘学生’,并根据学生在有标签数据上的表现迭代更新自身。我们通过基于目标掩码的无监督预训练阶段扩展了原模型,并对原模型、改进版本及多种独立基线进行了深入性能评估。实验在三种不同数据集、两种语言(英语和瑞典语)上进行。
原文摘要 · Abstract (English)
We extend and study a semi-supervised model for text classification proposed earlier by Hatefi et al. for classification tasks in which document classes are described by a small number of gold-labeled examples, while the majority of training examples is unlabeled. The model leverages the teacher-student architecture of Meta Pseudo Labels in which a ''teacher'' generates labels for originally unlabeled training data to train the ''student'' and updates its own model iteratively based on the performance of the student on the gold-labeled portion of the data. We extend the original model of Hatefi et al. by an unsupervised pre-training phase based on objective masking, and conduct in-depth performance evaluations of the original model, our extension, and various independent baselines. Experiments are performed using three different datasets in two different languages (English and Swedish).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。