arXiv:2410.00182cs.CLcs.AI2024-10被引 7

用大模型零样本识别灾难推文信息量与人道类别

Zero-Shot Classification of Crisis Tweets Using Instruction-Finetuned Large Language Models

  • 用指令微调大模型直接分类灾难推文,无需额外训练
  • 提供事件背景后,人道标签分类准确率显著提升
  • 不同数据集表现差异大,提示需警惕数据质量

社交媒体帖子常被视为灾害响应中的开源情报来源,预大模型时代自然语言处理技术已在灾难推文数据集上被评估。本文评估三种商用大语言模型(OpenAI GPT-4o、Gemini 1.5-flash-001 和 Anthropic Claude-3-5 Sonnet)在零样本分类短社交媒体帖子上的能力。在一个提示中,模型需完成两项任务:1)判断帖子是否具有人道主义语境下的信息性;2)对16种可能的人道主义类别进行排序并给出概率。待分类的帖子来自整合后的灾难推文数据集 CrisisBench。结果使用宏平均、加权和二值 F1 分数进行评估。信息性分类任务在无额外信息时表现较好,而提供推文采集时所发生的事件背景后,人道类别分类性能显著提升。此外,我们发现模型在不同数据集上的表现存在显著差异,这引发了对数据集质量的质疑。

原文摘要 · Abstract (English)

Social media posts are frequently identified as a valuable source of open-source intelligence for disaster response, and pre-LLM NLP techniques have been evaluated on datasets of crisis tweets. We assess three commercial large language models (OpenAI GPT-4o, Gemini 1.5-flash-001 and Anthropic Claude-3-5 Sonnet) capabilities in zero-shot classification of short social media posts. In one prompt, the models are asked to perform two classification tasks: 1) identify if the post is informative in a humanitarian context; and 2) rank and provide probabilities for the post in relation to 16 possible humanitarian classes. The posts being classified are from the consolidated crisis tweet dataset, CrisisBench. Results are evaluated using macro, weighted, and binary F1-scores. The informative classification task, generally performed better without extra information, while for the humanitarian label classification providing the event that occurred during which the tweet was mined, resulted in better performance. Further, we found that the models have significantly varying performance by dataset, which raises questions about dataset quality.

灾难推文零样本分类大模型应用人道分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。