用大模型文本嵌入分析监狱囚犯语言,发现同伴影响显著提升再犯预测准确率30%。
Recidivism and Peer Influence with LLM Text Embeddings in Low Security Correctional Facilities
- 将大模型文本嵌入与结构计量方法结合,构建语言同伴效应分析框架
- 文本特征使三年再犯预测准确率比传统变量高30%
- 适用于社会行为研究、刑事司法干预及大模型在人文社科中的应用
研究语言中的同伴效应至关重要,因其常反映影响经济结果的行为与人格特质。然而语言具有非结构化、非数值化和高维度的特点。本文将大语言模型(LLM)嵌入与结构计量识别方法结合,提出统一框架以识别语言中的同伴效应。该框架应用于低安全等级矫正机构中8万至12万份居民书面交流记录。基于LLM的语言画像比仅使用入院前协变量的预测方法,对三年内再犯的预测准确率提升30%,表明文本表示能捕捉有意义信号。在处理网络内生性问题的同时,分析多维语言嵌入的同伴效应。开发了适用于多变量结果、稀疏网络和多维潜变量的新型工具变量估计器。在现实稀疏条件下,方法实现根N一致性与渐近正态性,放宽了以往对稠密网络的假设。结果显示,居民语言特征存在显著同伴效应。
原文摘要 · Abstract (English)
Studying peer effects in language is critical because they often reflect behavioral and personality traits that are important determinants of economic outcomes. However, language is unstructured, non-numeric, and high-dimensional. We combine Large Language Model (LLM) embeddings with structural econometric identification to provide a unified framework for identifying peer effects in language. This unified framework is applied to 80,000-120,000 written exchanges among residents of low security correctional facilities. The LLM language profiles predict three-year recidivism 30\% more accurately than pre-entry covariates alone, showing that text representations capture meaningful signals. We analyze peer effects on multidimensional language embeddings while addressing network endogeneity. We develop novel instrumental variable estimators for peer effects that accommodate multivariate outcomes, sparse networks, and multidimensional latent variables. Our methods achieve root-N consistency and asymptotic normality under realistic sparsity conditions, relaxing the dense-network assumption. Results reveal significant peer effects in residents' language profiles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。