微调LLaMA模型可大幅提升情感分析性能,但零样本场景仍不理想。
LLaMA-Based Models for Aspect-Based Sentiment Analysis
- 用LLaMA系列模型微调做情感分析,提升整体表现。
- Orca-2在8个数据集上全面超越现有最优结果。
- 零样本/少样本时性能显著下降,适合有标注数据的场景。
尽管大语言模型(LLMs)在多种任务中展现出潜力,但在复杂方面级情感分析(ABSA)任务中的表现仍落后于微调模型。然而,针对ABSA任务进行微调的开源大模型潜力尚未被充分探索。本文聚焦于基于LLaMA的模型,评估其在四个任务和八个英文数据集上的表现,发现微调后的Orca-2模型在所有任务中均超越当前最佳结果。然而,所有模型在零样本和少样本场景下的表现均显著弱于全微调模型。此外,我们进行了错误分析,识别出微调模型面临的关键挑战。
原文摘要 · Abstract (English)
While large language models (LLMs) show promise for various tasks, their performance in compound aspect-based sentiment analysis (ABSA) tasks lags behind fine-tuned models. However, the potential of LLMs fine-tuned for ABSA remains unexplored. This paper examines the capabilities of open-source LLMs fine-tuned for ABSA, focusing on LLaMA-based models. We evaluate the performance across four tasks and eight English datasets, finding that the fine-tuned Orca~2 model surpasses state-of-the-art results in all tasks. However, all models struggle in zero-shot and few-shot scenarios compared to fully fine-tuned ones. Additionally, we conduct error analysis to identify challenges faced by fine-tuned models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。