评测大模型在各类灾害微博中的表现,发现其泛化能力强但处理洪水信息弱。
Evaluating Robustness of LLMs on Crisis-Related Microblogs across Events, Information Types, and Linguistic Features
- 对比六种大模型在真实灾害事件中的表现
- 洪水相关数据准确率低,少样本提示改善有限
- 私有模型优于开源模型,对错别字敏感
灾难期间,X(原推特)等微博客平台提供实时信息,但数据噪声大,需自动化筛选。传统监督学习模型泛化能力差,而大语言模型(LLMs)在自然语言理解上表现更优。本文系统评估了六种知名LLMs在多场真实灾害事件中处理灾情微博文本的表现。结果显示,尽管GPT-4o和GPT-4在跨事件、跨信息类型上具备更好泛化性,但多数模型在洪水类数据上表现不佳,即使提供示例(few-shot)也仅带来微弱提升,且难以识别紧急求助与需求等关键信息类别。我们进一步分析了不同语言特征对性能的影响,发现模型对拼写错误等语言变异敏感。最后,在零样本与少样本设置下对所有事件进行基准测试,发现专有模型整体优于开源模型。
原文摘要 · Abstract (English)
The widespread use of microblogging platforms like X (formerly Twitter) during disasters provides real-time information to governments and response authorities. However, the data from these platforms is often noisy, requiring automated methods to filter relevant information. Traditionally, supervised machine learning models have been used, but they lack generalizability. In contrast, Large Language Models (LLMs) show better capabilities in understanding and processing natural language out of the box. This paper provides a detailed analysis of the performance of six well-known LLMs in processing disaster-related social media data from a large-set of real-world events. Our findings indicate that while LLMs, particularly GPT-4o and GPT-4, offer better generalizability across different disasters and information types, most LLMs face challenges in processing flood-related data, show minimal improvement despite the provision of examples (i.e., shots), and struggle to identify critical information categories like urgent requests and needs. Additionally, we examine how various linguistic features affect model performance and highlight LLMs' vulnerabilities against certain features like typos. Lastly, we provide benchmarking results for all events across both zero- and few-shot settings and observe that proprietary models outperform open-source ones in all tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。