对比人类与ChatGPT在复杂社交媒体文本分类中的表现
A Comparison of Human and ChatGPT Classification Performance on Complex Social Media Data
- 测试GPT-3.5、GPT-4和GPT-4o在四种提示下的分类表现
- 模型对含细微语义的语言分类准确率不高,尤其在复杂语境中
- 适合研究者谨慎使用大模型进行需要语义理解的标注任务
生成式人工智能工具如ChatGPT在计算社会科学研究中日益普及。然而,对于其在复杂任务(如带有微妙语言特征的数据分类与标注)中的表现,仍需深入理解。本文评估GPT-4在一项此类任务中的表现,并与人工标注者对比。研究考察了ChatGPT的三个版本(3.5、4、4o),并设计四种提示风格,通过精确率、召回率和F1分数进行量化评估。定量与定性分析表明,尽管在提示中加入标签定义可能提升性能,但总体而言GPT-4在处理细微语义语言时仍存在困难。定性分析揭示四项具体发现,结果提示:在涉及复杂语义的分类任务中,使用ChatGPT应保持谨慎。
原文摘要 · Abstract (English)
Generative artificial intelligence tools, like ChatGPT, are an increasingly utilized resource among computational social scientists. Nevertheless, there remains space for improved understanding of the performance of ChatGPT in complex tasks such as classifying and annotating datasets containing nuanced language. Method. In this paper, we measure the performance of GPT-4 on one such task and compare results to human annotators. We investigate ChatGPT versions 3.5, 4, and 4o to examine performance given rapid changes in technological advancement of large language models. We craft four prompt styles as input and evaluate precision, recall, and F1 scores. Both quantitative and qualitative evaluations of results demonstrate that while including label definitions in prompts may help performance, overall GPT-4 has difficulty classifying nuanced language. Qualitative analysis reveals four specific findings. Our results suggest the use of ChatGPT in classification tasks involving nuanced language should be conducted with prudence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。