arXiv:2409.02370cs.CLcs.AI2024-09被引 5

测试大模型对情感的敏感度,发现其识别能力参差不齐。

Do Large Language Models Possess Sensitive to Sentiment?

  • 通过多任务实验评估主流大模型的情感识别与响应能力
  • 部分模型将强烈正向情绪误判为中性,且难以识别讽刺或反语
  • 模型表现差异显著,受架构与训练数据影响,适合关注情感理解的研究者

大型语言模型(LLMs)在语言理解方面展现出卓越能力,但如何全面评估其情感分析能力仍是挑战。本文研究了多个主流大模型在识别和回应文本情感(如正面、负面、中性)方面的表现。通过在多种情感基准数据集上进行实验,并与人工评价对比分析模型输出。结果表明,尽管模型具备一定情感敏感度,但在准确性和一致性上存在显著差异。例如,某些情况下模型会将强烈正向情感误判为中性,或无法识别文本中的讽刺与反语。此外,不同模型在同一数据集上的表现各异,这与模型架构和训练数据密切相关。研究强调需进一步优化训练过程以捕捉细微情感线索。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have recently displayed their extraordinary capabilities in language understanding. However, how to comprehensively assess the sentiment capabilities of LLMs continues to be a challenge. This paper investigates the ability of LLMs to detect and react to sentiment in text modal. As the integration of LLMs into diverse applications is on the rise, it becomes highly critical to comprehend their sensitivity to emotional tone, as it can influence the user experience and the efficacy of sentiment-driven tasks. We conduct a series of experiments to evaluate the performance of several prominent LLMs in identifying and responding appropriately to sentiments like positive, negative, and neutral emotions. The models' outputs are analyzed across various sentiment benchmarks, and their responses are compared with human evaluations. Our discoveries indicate that although LLMs show a basic sensitivity to sentiment, there are substantial variations in their accuracy and consistency, emphasizing the requirement for further enhancements in their training processes to better capture subtle emotional cues. Take an example in our findings, in some cases, the models might wrongly classify a strongly positive sentiment as neutral, or fail to recognize sarcasm or irony in the text. Such misclassifications highlight the complexity of sentiment analysis and the areas where the models need to be refined. Another aspect is that different LLMs might perform differently on the same set of data, depending on their architecture and training datasets. This variance calls for a more in-depth study of the factors that contribute to the performance differences and how they can be optimized.

情感分析大模型评测语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。