用大模型自动识别警民互动文本中的脆弱群体特征,准确且无明显偏见。
Using Instruction-Tuned Large Language Models to Identify Indicators of Vulnerability in Police Incident Narratives
- 用指令微调的大模型分析警方记录文本,识别心理健康、药物滥用等四类脆弱性
- 模型在无脆弱性场景下准确率高,可大幅减少人工标注工作量
- 测试显示模型对性别和种族的敏感度极低,结果稳定可靠
目的:比较指令微调的大语言模型(IT-LLMs)与人类编码员在分类警察-公众互动非结构化文本中是否存在脆弱性的表现。方法:基于波士顿警察局公开的警情叙述文本,向人类和IT-LLMs提供定性标注手册,对比其对四类脆弱性(心理疾病、物质滥用、酒精依赖、无家可归)的判断结果。研究探索了多种提示策略与模型规模,以及重复提示下的标签变异性。同时,采用反事实方法评估性别和种族这两个受保护特征对模型分类的影响。结果表明,IT-LLMs能有效辅助人工定性编码;尽管存在部分分歧,但模型在无脆弱性场景中表现优异,显著降低人工标注需求。反事实分析显示,改变文本中个体的性别或种族,对模型分类影响极小,接近随机水平。结论:IT-LLMs可高效辅助人类进行大规模自由文本分析,降低资源消耗,提升编码的精确性、透明度与可复现性。
原文摘要 · Abstract (English)
Objectives: Compare qualitative coding of instruction tuned large language models (IT-LLMs) against human coders in classifying the presence or absence of vulnerability in routinely collected unstructured text that describes police-public interactions. Evaluate potential bias in IT-LLM codings. Methods: Analyzing publicly available text narratives of police-public interactions recorded by Boston Police Department, we provide humans and IT-LLMs with qualitative labelling codebooks and compare labels generated by both, seeking to identify situations associated with (i) mental ill health; (ii) substance misuse; (iii) alcohol dependence; and (iv) homelessness. We explore multiple prompting strategies and model sizes, and the variability of labels generated by repeated prompts. Additionally, to explore model bias, we utilize counterfactual methods to assess the impact of two protected characteristics - race and gender - on IT-LLM classification. Results: Results demonstrate that IT-LLMs can effectively support human qualitative coding of police incident narratives. While there is some disagreement between LLM and human generated labels, IT-LLMs are highly effective at screening narratives where no vulnerabilities are present, potentially vastly reducing the requirement for human coding. Counterfactual analyses demonstrate that manipulations to both gender and race of individuals described in narratives have very limited effects on IT-LLM classifications beyond those expected by chance. Conclusions: IT-LLMs offer effective means to augment human qualitative coding in a way that requires much lower levels of resource to analyze large unstructured datasets. Moreover, they encourage specificity in qualitative coding, promote transparency, and provide the opportunity for more standardized, replicable approaches to analyzing large free-text police data sources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。