对比大模型与人工在情感、立场等分析中的表现,发现大模型稳定可靠。
Evaluating Large Language Models Against Human Annotators in Latent Content Analysis: Sentiment, Political Leaning, Emotional Intensity, and Sarcasm
- 用7个主流大模型和33名人工标注者对比分析文本情感、立场等四维度。
- 大模型在情感和立场判断上可靠性高,情绪强度一致性优于人类。
- 大模型时间稳定性极佳,适合长期内容分析,但讽刺识别仍不足。
在数字通信快速发展的时代,海量文本数据亟需高效的隐含内容分析方法以提取有意义的洞察。大语言模型(LLMs)为自动化这一过程提供了可能,但缺乏对它们在多个维度上与人类标注者表现的全面评估。本研究评估了七种前沿大模型(包括GPT-4、Gemini、Llama、Mixtral等变体)在情感、政治倾向、情绪强度和反讽检测方面的可靠性、一致性和质量,相较于33名人类标注者。共100个精选文本样本由33名标注者完成3,300次人工标注,8种大模型版本生成19,200次标注,并在三个时间点进行评估以考察时序一致性。采用Krippendorff's alpha衡量评分者间可靠性,组内相关系数评估时间一致性。结果表明,人类与大模型在情感分析和政治倾向判断中均表现出高可靠性,且大模型内部一致性高于人类;在情绪强度方面,大模型一致性更高,但人类评分显著更高;两者在反讽检测中均表现不佳,一致性低。大模型在所有维度上均展现出优异的时间一致性,表明其性能稳定。研究结论:大模型(尤其是GPT-4)可在情感与政治倾向分析中有效复现人类判断,但情绪强度解读仍需人类经验。该研究展示了大模型在特定领域隐含内容分析中具备持续高质量表现的潜力。
原文摘要 · Abstract (English)
In the era of rapid digital communication, vast amounts of textual data are generated daily, demanding efficient methods for latent content analysis to extract meaningful insights. Large Language Models (LLMs) offer potential for automating this process, yet comprehensive assessments comparing their performance to human annotators across multiple dimensions are lacking. This study evaluates the reliability, consistency, and quality of seven state-of-the-art LLMs, including variants of OpenAI's GPT-4, Gemini, Llama, and Mixtral, relative to human annotators in analyzing sentiment, political leaning, emotional intensity, and sarcasm detection. A total of 33 human annotators and eight LLM variants assessed 100 curated textual items, generating 3,300 human and 19,200 LLM annotations, with LLMs evaluated across three time points to examine temporal consistency. Inter-rater reliability was measured using Krippendorff's alpha, and intra-class correlation coefficients assessed consistency over time. The results reveal that both humans and LLMs exhibit high reliability in sentiment analysis and political leaning assessments, with LLMs demonstrating higher internal consistency than humans. In emotional intensity, LLMs displayed higher agreement compared to humans, though humans rated emotional intensity significantly higher. Both groups struggled with sarcasm detection, evidenced by low agreement. LLMs showed excellent temporal consistency across all dimensions, indicating stable performance over time. This research concludes that LLMs, especially GPT-4, can effectively replicate human analysis in sentiment and political leaning, although human expertise remains essential for emotional intensity interpretation. The findings demonstrate the potential of LLMs for consistent and high-quality performance in certain areas of latent content analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。