系统评估40个暴力预测模型,发现多数存在偏差且临床用处有限。
Machine learning for violence prediction: a systematic review and critical appraisal
- 系统检索38项研究,整合40个机器学习模型的性能数据
- 仅8项报告校准性能,3项完成外部验证,多数模型易过拟合
- 强调需提升模型透明度与可解释性,适合临床风险监控场景
目的:系统回顾机器学习模型在预测暴力行为中的应用,综合评估其有效性、实用性及表现。方法:截至2025年9月,系统检索九个文献数据库及谷歌学术,筛选开发和/或验证类研究。通过汇总判别力与校准性能指标,并评估偏倚风险与临床实用性,对研究质量进行评价。结果:共识别38项研究,报告了40个模型的开发与验证。多数研究使用受试者工作特征曲线下面积(AUC)作为判别力指标,范围为0.68–0.99。仅有8项报告校准性能,3项完成外部验证。31项研究存在高偏倚风险,主要集中在分析环节;仅3项为低风险。整体临床实用性差,原因包括样本量小导致过拟合、报告不透明、泛化能力弱。结论:尽管当前黑箱模型在临床中应用受限,但可能有助于识别高风险个体。建议五个关键改进方向:(i) 提升方法学质量与跨学科合作;(ii) 仅在复杂数据中使用黑箱算法;(iii) 引入动态预测实现风险持续监测;(iv) 采用可解释方法提升可信度;(v) 在合适场景下应用因果机器学习。
原文摘要 · Abstract (English)
Purpose To conduct a systematic review of machine learning models for predicting violent behaviour by synthesising and appraising their validity, usefulness, and performance. Methods We systematically searched nine bibliographic databases and Google Scholar up to September 2025 for development and/or validation studies on machine learning methods for predicting all forms of violent behaviour. We synthesised the results by summarising discrimination and calibration performance statistics and evaluated study quality by examining risk of bias and clinical utility. Results We identified 38 studies reporting the development and validation of 40 models. Most studies reported Area Under the Curve (AUC) as the discrimination statistic with a range of 0.68-0.99. Only eight studies reported calibration performance, and three studies reported external validation. 31 studies had a high risk of bias, mainly in the analysis domain, and three studies had low risk of bias. The overall clinical utility of violence prediction models is poor, as indicated by risks of overfitting due to small samples, lack of transparent reporting, and low generalisability. Conclusion Although black box machine learning models currently have limited applicability in clinical settings, they may show promise for identifying high-risk individuals. We recommend five key considerations for violence prediction modelling: (i) ensuring methodological quality (e.g. following guidelines) and interdisciplinary collaborations; (ii) using black box algorithms only for highly complex data; (iii) incorporating dynamic predictions to allow for risk monitoring; (iv) developing more trustworthy algorithms using explainable methods; and (v) applying causal machine learning approaches where appropriate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。