arXiv:2509.04982cs.CLcs.IR2025-09中稿 · LDD@ECAI 2025被引 1

小模型在短文本情感分析中,数据增强有效但持续预训练易引入噪声。

Optimizing Small Transformer-Based Language Models for Multi-Label Sentiment Analysis in Short Texts

  • 用生成式数据增强提升小模型性能
  • 在增广数据上继续预训练反而降低准确率
  • 修改分类头效果有限,适合资源少的场景

短文本情感分类面临类别不平衡、训练样本少及情感标签主观性强等问题,短文本的上下文有限进一步加剧了歧义和数据稀疏性。本文评估了小型Transformer模型(如BERT和RoBERTa,参数少于10亿)在多标签情感分类任务中的表现,重点关注短文本场景。研究考察了三个关键因素:(1) 领域特定的持续预训练,(2) 使用自动生成样本的数据增强,特别是生成式数据增强,(3) 分类头的结构变化。实验结果表明,数据增强能提升分类性能,但在增广数据上进行持续预训练反而引入噪声,导致准确率下降;同时,分类头结构调整带来的收益微乎其微。这些发现为资源受限环境下优化BERT类模型提供了实用指导,并有助于改进短文本情感分类策略。

原文摘要 · Abstract (English)

Sentiment classification in short text datasets faces significant challenges such as class imbalance, limited training samples, and the inherent subjectivity of sentiment labels -- issues that are further intensified by the limited context in short texts. These factors make it difficult to resolve ambiguity and exacerbate data sparsity, hindering effective learning. In this paper, we evaluate the effectiveness of small Transformer-based models (i.e., BERT and RoBERTa, with fewer than 1 billion parameters) for multi-label sentiment classification, with a particular focus on short-text settings. Specifically, we evaluated three key factors influencing model performance: (1) continued domain-specific pre-training, (2) data augmentation using automatically generated examples, specifically generative data augmentation, and (3) architectural variations of the classification head. Our experiment results show that data augmentation improves classification performance, while continued pre-training on augmented datasets can introduce noise rather than boost accuracy. Furthermore, we confirm that modifications to the classification head yield only marginal benefits. These findings provide practical guidance for optimizing BERT-based models in resource-constrained settings and refining strategies for sentiment classification in short-text datasets.

小模型情感分析数据增强短文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。