用大模型从海量文本中自动提取枪击案关键信息,助力司法调查。
Knowledge Acquisition on Mass-shooting Events via LLMs for AI-Driven Justice
- 基于大模型少样本提示,自动识别枪击案中的嫌疑人、受害者等实体。
- GPT-4o在真实数据集上表现最优,微平均准确率、召回率和F1值均最高。
- 适用于法律、公安及政策研究者,提升重大事件信息处理效率。
大规模枪击事件对公共安全构成严峻挑战,产生大量非结构化文本数据,阻碍有效调查与公共政策制定。尽管需求迫切,但以往研究极少能自动化提取此类事件的关键信息以支持法律与侦查工作。本文首次构建面向大规模枪击事件的知识获取数据集,采用命名实体识别(NER)技术,聚焦识别嫌疑人、受害者、地点及作案工具等关键实体。该过程依托大语言模型(LLM)的少样本提示机制,从新闻报道、警方报告及社交媒体等多种来源高效提取并组织关键信息。在真实世界枪击事件语料上的实验表明,GPT-4o在微平均精确率、召回率和F1值上均表现最佳;o1-mini则展现出良好性能,适合作为资源受限场景下的替代方案。同时发现,增加提示样本数量可提升所有模型表现,尤其对GPT-4o和o1-mini效果更显著,凸显其在少样本学习中的强适应性。
原文摘要 · Abstract (English)
Mass-shooting events pose a significant challenge to public safety, generating large volumes of unstructured textual data that hinder effective investigations and the formulation of public policy. Despite the urgency, few prior studies have effectively automated the extraction of key information from these events to support legal and investigative efforts. This paper presented the first dataset designed for knowledge acquisition on mass-shooting events through the application of named entity recognition (NER) techniques. It focuses on identifying key entities such as offenders, victims, locations, and criminal instruments, that are vital for legal and investigative purposes. The NER process is powered by Large Language Models (LLMs) using few-shot prompting, facilitating the efficient extraction and organization of critical information from diverse sources, including news articles, police reports, and social media. Experimental results on real-world mass-shooting corpora demonstrate that GPT-4o is the most effective model for mass-shooting NER, achieving the highest Micro Precision, Micro Recall, and Micro F1-scores. Meanwhile, o1-mini delivers competitive performance, making it a resource-efficient alternative for less complex NER tasks. It is also observed that increasing the shot count enhances the performance of all models, but the gains are more substantial for GPT-4o and o1-mini, highlighting their superior adaptability to few-shot learning scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。