对比四款大模型在波斯语社交文本情感与情绪分析中的表现
A Comparative Evaluation of Large Language Models for Persian Sentiment Analysis and Emotion Detection in Social Media Texts
- 统一提示与参数,公平比较四款主流大模型性能
- 情绪检测任务难度高于情感分析,所有模型准确率均超85%
- GPT-4o略胜于精度,Gemini 2.0 Flash最经济高效
本研究对四款先进大语言模型(Claude 3.7 Sonnet、DeepSeek-V3、Gemini 2.0 Flash、GPT-4o)在波斯语社交媒体文本的情感分析与情绪检测任务中进行了全面对比评估。采用平衡的波斯语数据集,包含900条情感分析样本(正面、负面、中性)和1800条情绪检测样本(愤怒、恐惧、快乐、仇恨、悲伤、惊讶)。通过一致提示、统一处理参数及精确度、召回率、F1分数等指标进行严格比较,并分析误分类模式。结果显示,所有模型均达到可接受性能水平;三款最优模型间无显著差异。其中,GPT-4o在两项任务中均取得最高原始准确率,而Gemini 2.0 Flash在成本效率上表现最佳。情绪检测任务整体难度高于情感分析,且误分类模式反映出波斯语特有的语言与文化挑战。研究为波斯语自然语言处理提供了性能基准,为多语言AI系统部署中的模型选型提供依据。
原文摘要 · Abstract (English)
This study presents a comprehensive comparative evaluation of four state-of-the-art Large Language Models (LLMs)--Claude 3.7 Sonnet, DeepSeek-V3, Gemini 2.0 Flash, and GPT-4o--for sentiment analysis and emotion detection in Persian social media texts. Comparative analysis among LLMs has witnessed a significant rise in recent years, however, most of these analyses have been conducted on English language tasks, creating gaps in understanding cross-linguistic performance patterns. This research addresses these gaps through rigorous experimental design using balanced Persian datasets containing 900 texts for sentiment analysis (positive, negative, neutral) and 1,800 texts for emotion detection (anger, fear, happiness, hate, sadness, surprise). The main focus was to allow for a direct and fair comparison among different models, by using consistent prompts, uniform processing parameters, and by analyzing the performance metrics such as precision, recall, F1-scores, along with misclassification patterns. The results show that all models reach an acceptable level of performance, and a statistical comparison of the best three models indicates no significant differences among them. However, GPT-4o demonstrated a marginally higher raw accuracy value for both tasks, while Gemini 2.0 Flash proved to be the most cost-efficient. The findings indicate that the emotion detection task is more challenging for all models compared to the sentiment analysis task, and the misclassification patterns can represent some challenges in Persian language texts. These findings establish performance benchmarks for Persian NLP applications and offer practical guidance for model selection based on accuracy, efficiency, and cost considerations, while revealing cultural and linguistic challenges that require consideration in multilingual AI system deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。