用大模型分析20万篇西语报纸,重建哥伦比亚冲突历史
Using LLMs to create analytical datasets: A case study of reconstructing the historical memory of Colombia
- 用GPT解析20多万篇西语暴力相关报纸文章
- 发现暴力与古柯作物清除政策存在关联性
- 为政策研究提供可扩展的文本分析新范式
哥伦比亚长期处于武装冲突中,但政府过去未将系统记录暴力事件列为重点。这导致公开冲突信息匮乏,历史记忆缺失。本研究利用GPT这一大型语言模型,对超过20万篇西班牙语暴力相关报纸文章进行阅读与问答处理,构建分析数据集。基于该数据集,开展描述性分析及暴力与古柯作物清除政策关系的研究,展示此类数据可支持的政策分析潜力。研究证明,大语言模型使对大规模文本语料进行以往无法实现的深度考察成为可能。
原文摘要 · Abstract (English)
Colombia has been submerged in decades of armed conflict, yet until recently, the systematic documentation of violence was not a priority for the Colombian government. This has resulted in a lack of publicly available conflict information and, consequently, a lack of historical accounts. This study contributes to Colombia's historical memory by utilizing GPT, a large language model (LLM), to read and answer questions about over 200,000 violence-related newspaper articles in Spanish. We use the resulting dataset to conduct both descriptive analysis and a study of the relationship between violence and the eradication of coca crops, offering an example of policy analyses that such data can support. Our study demonstrates how LLMs have opened new research opportunities by enabling examinations of large text corpora at a previously infeasible depth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。