ConfliBERT专用于识别政治冲突文本中的行为与主体,比通用大模型更准更快。
ConfliBERT: A Language Model for Political Conflict
- 基于新闻和冲突数据训练,专注提取政治暴力事件中的主体与行为
- 在BBC、re3d、GTD数据上准确率、召回率均优于Gemma、Llama、Qwen等模型
- 推理速度比通用大模型快数百倍,适合大规模冲突文本分析
冲突研究者传统上使用规则方法从新闻报道中提取政治暴力信息。近年来自然语言处理技术已超越固定规则方法。我们回顾了近期提出的ConfliBERT语言模型(Hu et al. 2022),该模型可处理政治与暴力相关文本,用于提取其中的行动者与行为分类。经微调后,在其相关领域内,ConfliBERT在准确率、精确率和召回率上均优于Google的Gemma 2 (9B)、Meta的Llama 3.1 (7B)以及阿里巴巴的Qwen 2.5 (14B)等大型语言模型。同时,其推理速度比这些通用大模型快数百倍。实验基于BBC、re3d及全球恐怖主义数据库(GTD)的文本数据进行验证。
原文摘要 · Abstract (English)
Conflict scholars have used rule-based approaches to extract information about political violence from news reports and texts. Recent Natural Language Processing developments move beyond rigid rule-based approaches. We review our recent ConfliBERT language model (Hu et al. 2022) to process political and violence related texts. The model can be used to extract actor and action classifications from texts about political conflict. When fine-tuned, results show that ConfliBERT has superior performance in accuracy, precision and recall over other large language models (LLM) like Google's Gemma 2 (9B), Meta's Llama 3.1 (7B), and Alibaba's Qwen 2.5 (14B) within its relevant domains. It is also hundreds of times faster than these more generalist LLMs. These results are illustrated using texts from the BBC, re3d, and the Global Terrorism Dataset (GTD).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。