GPT-4 Turbo可准确模拟专家对烟草产品社交媒体情绪的判断,达80%以上。
Large Language Models' Accuracy in Emulating Human Experts' Evaluation of Public Sentiments about Heated Tobacco Products on Social Media
- 用GPT-4 Turbo多次生成答案取多数投票,提升情绪分类准确率。
- 在推特和脸书上对加热烟草产品的判断准确率达77%~81.7%。
- 适合需快速分析公众情绪的研究者,但需注意中性情绪易误判。
社交媒体上对替代烟草产品的态度分析对控烟研究至关重要。大型语言模型(LLMs)可缓解人工情绪分析的耗时问题。本研究评估了GPT-3.5与GPT-4 Turbo在复制人类专家对加热烟草产品(HTPs)相关社交消息情绪判断上的准确性。使用GPT-3.5和GPT-4 Turbo对500条脸书与500条推特消息进行分类,包含反HTP、支持HTP及中立三类。每条消息由模型重复评估最多20次,取多数标签与人工标注对比。结果显示,GPT-3.5在脸书消息上准确率为61.2%,推特为57.0%;GPT-4 Turbo表现更优,脸书达81.7%,推特为77.0%。仅用三次响应实例,GPT-4 Turbo即达到20次的99%准确率。其在反/正向消息上的准确率高于中立类。GPT-3.5常将反/正向消息误判为中立或无关,而GPT-4 Turbo在各类别均显著改善。结论:大模型可用于HTP相关社交情绪分析,其中GPT-4 Turbo准确率接近人类专家,约80%,但不同情绪类别间差异可能导致整体情绪误判风险。
原文摘要 · Abstract (English)
Sentiment analysis of alternative tobacco products on social media is important for tobacco control research. Large Language Models (LLMs) can help streamline the labor-intensive human sentiment analysis process. This study examined the accuracy of LLMs in replicating human sentiment evaluation of social media messages about heated tobacco products (HTPs). The research used GPT-3.5 and GPT-4 Turbo to classify 500 Facebook and 500 Twitter messages, including anti-HTPs, pro-HTPs, and neutral messages. The models evaluated each message up to 20 times, and their majority label was compared to human evaluators. Results showed that GPT-3.5 accurately replicated human sentiment 61.2% of the time for Facebook messages and 57.0% for Twitter messages. GPT-4 Turbo performed better, with 81.7% accuracy for Facebook and 77.0% for Twitter. Using three response instances, GPT-4 Turbo achieved 99% of the accuracy of twenty instances. GPT-4 Turbo also had higher accuracy for anti- and pro-HTPs messages compared to neutral ones. Misclassifications by GPT-3.5 often involved anti- or pro-HTPs messages being labeled as neutral or irrelevant, while GPT-4 Turbo showed improvements across all categories. In conclusion, LLMs can be used for sentiment analysis of HTP-related social media messages, with GPT-4 Turbo reaching around 80% accuracy compared to human experts. However, there's a risk of misrepresenting overall sentiment due to differences in accuracy across sentiment categories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。