arXiv:2506.07148cs.CL2025-06

用结构化提示增强数据,保持语义不变并提升情感分类准确率

Semantic-preserved Augmentation with Confidence-weighted Fine-tuning for Aspect Category Sentiment Analysis

  • 用结构化模板让大模型生成语义一致的增强数据
  • 在4个基准数据集上超越现有方法,性能最优
  • 适合低资源场景下情感分析任务的模型优化

大型语言模型(LLM)是解决低资源场景下数据稀缺问题的有效方法。现有研究多采用手工设计的提示引导LLM进行数据增强。本文提出一种面向方面类别情感分析(ACSA)的数据增强策略,通过提供结构化提示模板,使LLM生成预设内容,同时保留原句语义并提升语言多样性。此外,采用后处理技术进一步确保生成句与原始句的语义一致性。增强数据扩大了训练分布的语义覆盖范围,帮助模型更好理解方面类别与情感极性之间的关系,提升推理能力。此外,提出置信度加权微调策略,促使模型生成更自信、更准确的情感极性预测。相比强大且最新的方法,本方法在四个基准数据集上始终优于所有基线。

原文摘要 · Abstract (English)

Large language model (LLM) is an effective approach to addressing data scarcity in low-resource scenarios. Recent existing research designs hand-crafted prompts to guide LLM for data augmentation. We introduce a data augmentation strategy for the aspect category sentiment analysis (ACSA) task that preserves the original sentence semantics and has linguistic diversity, specifically by providing a structured prompt template for an LLM to generate predefined content. In addition, we employ a post-processing technique to further ensure semantic consistency between the generated sentence and the original sentence. The augmented data increases the semantic coverage of the training distribution, enabling the model better to understand the relationship between aspect categories and sentiment polarities, enhancing its inference capabilities. Furthermore, we propose a confidence-weighted fine-tuning strategy to encourage the model to generate more confident and accurate sentiment polarity predictions. Compared with powerful and recent works, our method consistently achieves the best performance on four benchmark datasets over all baselines.

情感分析数据增强大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。