arXiv:2505.06145cs.CLcs.LG2025-05被引 7

用双损失策略提升Transformer在少样本文本分类中的表现

Towards Robust Few-Shot Text Classification Using Transformer Architectures and Dual Loss Strategies

  • 结合自适应微调与对比学习,优化模型特征提取能力
  • 5次样本下准确率显著提升,缓解了过拟合问题
  • 适合低资源场景下的文本分类任务研究者参考

少样本文本分类在低资源环境下具有重要应用价值。本文提出一种结合自适应微调、对比学习和正则化优化的策略,以提升基于Transformer模型的分类性能。在FewRel 2.0数据集上的实验表明,T5-small、DeBERTa-v3和RoBERTa-base在少样本任务中表现良好,尤其在5次样本设置下能更有效捕捉文本特征并提高分类准确率。实验还发现不同关系类别的分类难度差异显著,部分类别语义边界模糊或特征分布复杂,标准交叉熵损失难以学习判别性信息。通过引入对比损失和正则化损失,增强了模型泛化能力,有效缓解了少样本环境下的过拟合问题。此外,研究结果表明,具备更强自注意力机制的Transformer或生成式架构有助于提升少样本分类的稳定性和准确性。

原文摘要 · Abstract (English)

Few-shot text classification has important application value in low-resource environments. This paper proposes a strategy that combines adaptive fine-tuning, contrastive learning, and regularization optimization to improve the classification performance of Transformer-based models. Experiments on the FewRel 2.0 dataset show that T5-small, DeBERTa-v3, and RoBERTa-base perform well in few-shot tasks, especially in the 5-shot setting, which can more effectively capture text features and improve classification accuracy. The experiment also found that there are significant differences in the classification difficulty of different relationship categories. Some categories have fuzzy semantic boundaries or complex feature distributions, making it difficult for the standard cross entropy loss to learn the discriminative information required to distinguish categories. By introducing contrastive loss and regularization loss, the generalization ability of the model is enhanced, effectively alleviating the overfitting problem in few-shot environments. In addition, the research results show that the use of Transformer models or generative architectures with stronger self-attention mechanisms can help improve the stability and accuracy of few-shot classification.

少样本分类Transformer对比学习文本分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。