用微调大模型分析英国警情记录,识别弱势群体指标。
Using Fine-Tuned LLMs to Identify Indicators of Vulnerability in UK Police Incident Logs

- 用多阶段流程结合大模型推理与人工校正,处理警方非结构化数据。
- 约五分之一警情涉及心理健康问题,其他指标更少,但模型易高估。
- 需大量人工干预和统计修正,结果仅适合宏观分析,不适合个体决策。
目的:理解常规警务中有多少涉及弱势人群,可为资源分配、培训和跨部门协作提供依据,但行政数据信息有限。本文探讨基于开源美国警情数据训练的LLM分类流程,是否能适配用于估算英国警情叙述中四种弱势指标(心理健康问题、药物滥用、酒精依赖、无家可归)的出现频率,并判断其输出能否作为可信测量。方法:分析近3,000条去标识化的英国某警力部门警情记录,采用多阶段流程,包括重复模型推理、标签聚合、结构化人工审查及统计校正。整个流程在本地部署的开源权重LLM上运行,符合警方对数据安全的要求。结果:大模型可在大规模下生成有意义但不完美的估计。心理健康问题出现在约五分之一的警情中,其他指标频率更低。然而,直接使用模型不可靠:单次推理结果不稳定,聚合输出系统性高估指标,相比人工判断。纠正这些偏差需大量人工投入与统计调整,仍存在显著不确定性。结论:尽管大模型可从非结构化警情数据中提取信息,但其输出无法在未经严谨方法支持的情况下视为有效度量。在人口层面,可获得可接受的估计,但成本高昂;在个体层面,错误频繁且不可预测,不适合用于操作决策。本研究揭示了大模型在实际应用中的潜力与局限。
原文摘要 · Abstract (English)
Purpose: Understanding how much of routine policing involves vulnerable people could inform resourcing, training, and multi-agency response, yet administrative data provide limited insight. We explore whether an LLM-based classification pipeline, developed on open-source US police data, can be adapted to estimate the prevalence of four vulnerability indicators - mental ill health, substance misuse, alcohol dependence, and homelessness - in UK police incident narratives, and when outputs can be treated as defensible measurements. Methods: We analyse nearly 3,000 de-identified incident logs from a UK police force, using a multi-stage pipeline combining repeated model inference, label aggregation, structured human review, and statistical correction. The pipeline runs on a locally hosted open-weight LLM, reflecting the secure environments police must work in. Results: LLMs can produce meaningful, if imperfect, prevalence estimates at scale. Mental ill health indicators are present in approximately one in five incidents, with lower prevalence for other indicators. However, naive LLM deployment is unreliable: single-pass classifications are unstable, and aggregated outputs systematically over-assign indicators relative to human judgement. Correcting these biases required substantial human input and statistical adjustment, leaving considerable uncertainty. Conclusions: While LLMs can extract information from unstructured police data, their outputs cannot be treated as valid measurements without careful methodological support. At the population level, defensible estimates are achievable but resource-intensive; at the individual level, errors remain frequent and unpredictable, limiting suitability for operational decisions. This study highlights both the potential and the constraints of LLM-based measurement in applied settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。