arXiv:2412.02279cs.CLcs.AI2024-12被引 14

大模型在情感分析中表现超越小模型,无需微调也能高效完成任务。

A Comprehensive Evaluation of Large Language Models on Aspect-Based Sentiment Analysis

  • 统一任务框架,覆盖8个子任务、13个数据集,评估6种大模型。
  • 微调范式下,大模型性能超过微调小模型;零样本范式下仍具竞争力。
  • 提出三种演示选择策略,提升少样本学习效果,适合工业级应用者参考。

近年来,大语言模型(LLMs)在自然语言处理领域备受关注,凭借强大的推理与生成能力,革新了众多下游任务。例如,上下文学习(ICL)提供免微调范式,使大模型通过类比学习执行下游任务;在有充足训练数据的微调范式中,参数高效微调(PEFT)以低成本实现接近全量微调的性能。然而,这些技术在方面级情感分析(ABSA)领域的应用尚未充分探索。此前研究仅用随机输入输出对作为ICL演示,导致评估不完整且浅显。本文首次开展全面评估,涵盖13个数据集、8个ABSA子任务和6种大模型。我们设计统一任务框架,整合多模型、多子任务、多范式。针对微调依赖范式,采用基于指令的多任务学习高效微调;针对免微调范式,提出三种演示选择策略以激发少样本能力。大量实验表明,在微调范式下,大模型性能超越微调小模型(SLMs);更重要的是,在免微调范式中,小模型失效时,大模型结合ICL仍展现出显著潜力,甚至在部分子任务上媲美微调小模型。

原文摘要 · Abstract (English)

Recently, Large Language Models (LLMs) have garnered increasing attention in the field of natural language processing, revolutionizing numerous downstream tasks with powerful reasoning and generation abilities. For example, In-Context Learning (ICL) introduces a fine-tuning-free paradigm, allowing out-of-the-box LLMs to execute downstream tasks by analogy learning without any fine-tuning. Besides, in a fine-tuning-dependent paradigm where substantial training data exists, Parameter-Efficient Fine-Tuning (PEFT), as the cost-effective methods, enable LLMs to achieve excellent performance comparable to full fine-tuning. However, these fascinating techniques employed by LLMs have not been fully exploited in the ABSA field. Previous works probe LLMs in ABSA by merely using randomly selected input-output pairs as demonstrations in ICL, resulting in an incomplete and superficial evaluation. In this paper, we shed light on a comprehensive evaluation of LLMs in the ABSA field, involving 13 datasets, 8 ABSA subtasks, and 6 LLMs. Specifically, we design a unified task formulation to unify ``multiple LLMs for multiple ABSA subtasks in multiple paradigms.'' For the fine-tuning-dependent paradigm, we efficiently fine-tune LLMs using instruction-based multi-task learning. For the fine-tuning-free paradigm, we propose 3 demonstration selection strategies to stimulate the few-shot abilities of LLMs. Our extensive experiments demonstrate that LLMs achieve a new state-of-the-art performance compared to fine-tuned Small Language Models (SLMs) in the fine-tuning-dependent paradigm. More importantly, in the fine-tuning-free paradigm where SLMs are ineffective, LLMs with ICL still showcase impressive potential and even compete with fine-tuned SLMs on some ABSA subtasks.

大模型情感分析零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。