小模型可精准识别儿童福利记录中的5类成瘾物质,助力临床决策。
Validation of a Small Language Model for DSM-5 Substance Category Classification in Child Welfare Records
- 用200亿参数本地模型分两阶段识别7类物质,基于DSM-5标准
- 酒精、大麻等5类物质分类准确率92%~100%,一致性达0.94以上
- 适合需隐私保护的本地部署场景,如政府机构与医疗机构
背景:近期研究显示大型语言模型(LLMs)可在儿童福利叙述中完成二元分类任务,检测成瘾问题、家庭暴力和枪支涉及等概念。但小型、可本地部署的模型能否超越二元检测,从这些文本中识别具体物质类型尚无验证。目标:验证一个本地部署的大型语言模型分类器,用于识别与DSM-5分类一致的特定物质类型。方法:采用一个本地部署的200亿参数语言模型,对美国中西部某州的儿童虐待调查叙述进行分类。此前被标记为含物质相关问题的记录进入第二阶段,针对7个DSM-5物质类别进行分类。通过专家人工审查900个分层样本,评估分类精度、召回率及方法间一致性(Cohen's kappa)。使用约15000条独立分类记录评估测试-重测稳定性。结果:五类物质(酒精、大麻、阿片类、兴奋剂、镇静剂/催眠剂/抗焦虑药)达到几乎完美的方法间一致性(kappa = 0.94–1.00),分类精度为92%至100%。两类低发生率类别(致幻剂、吸入剂)表现较差。测试-重测一致性在7类中为92.1%至99.1%。结论:小型本地部署语言模型能可靠地从儿童福利行政文本中分类特定物质类型,将先前的二元分类工作拓展至多标签物质识别。
原文摘要 · Abstract (English)
Background: Recent studies have demonstrated that large language models (LLMs) can perform binary classification tasks on child welfare narratives, detecting the presence or absence of constructs such as substance-related problems, domestic violence, and firearms involvement. Whether smaller, locally deployable models can move beyond binary detection to classify specific substance types from these narratives remains untested. Objective: To validate a locally hosted LLM classifier for identifying specific substance types aligned with DSM-5 categories in child welfare investigation narratives. Methods: A locally hosted 20-billion-parameter LLM classified child maltreatment investigation narratives from a Midwestern U.S. state. Records previously identified as containing substance-related problems were passed to a second classification stage targeting seven DSM-5 substance categories. Expert human review of 900 stratified cases assessed classification precision, recall, and inter-method reliability (Cohen's kappa). Test-retest stability was evaluated using approximately 15,000 independently classified records. Results: Five substance categories achieved almost perfect inter-method agreement (kappa = 0.94-1.00): alcohol, cannabis, opioid, stimulant, and sedative/hypnotic/anxiolytic. Classification precision ranged from 92% to 100% for these categories. Two low-prevalence categories (hallucinogen, inhalant) performed poorly. Test-retest agreement ranged from 92.1% to 99.1% across the seven categories. Conclusions: A small, locally hosted LLM can reliably classify substance types from child welfare administrative text, extending prior work on binary classification to multi-label substance identification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。