用大模型分析儿童保护报告,自动评估父母配合度。
Reasoning Language Models for complex assessments tasks: Evaluating parental cooperation from child protection case reports
- 用推理型大模型分析案例报告中的父母配合行为。
- 最大模型准确率达89%,对母亲评估更准(93%)。
- 适合研究儿童保护中性别偏见的学者或从业者。
目的:推理语言模型(RLMs)在复杂推理任务中表现优异。本文考察其在使用案例报告评估儿童保护干预中父母配合度的潜力,该因素常因信息模糊且矛盾而难以判断。方法:构建四阶段工作流:(1)收集案例报告,(2)基于推理评估父母配合度,(3)自动提取类别,(4)案例标注。对比不同参数规模的RLMs(255B、32B、4B)与人工验证数据的表现。两名专家评审员独立分类加权随机样本报告。结果:最大规模模型准确率达89%,优于初始方法(80%)。对母亲的评估准确率(93%)高于父亲(85%),专家评审也呈现相似差异。结论:RLMs的推理能力可有效评估父母配合等复杂案情因素。对父亲评估准确率较低,支持了儿童保护领域对母亲存在更强专业关注的观点。
原文摘要 · Abstract (English)
Purpose: Reasoning language models (RLMs) have demonstrated significant advances in solving complex reasoning tasks. We examined their potential to assess parental cooperation during CPS interventions using case reports, a case factor characterized by ambiguous and conflicting information. Methods: A four stage workflow comprising (1) case reports collection, (2) reasoning-based assessment of parental cooperation, (3) automated category extraction, and (4) case labeling was developed. The performance of RLMs with different parameter sizes (255B, 32B, 4B) was compared against human validated data. Two expert human reviewers (EHRs) independently classified a weighted random sample of reports. Results: The largest RLM achieved the highest accuracy (89%), outperforming the initial approach (80%). Classification accuracy was higher for mothers (93%) than for fathers (85%), and EHRs exhibited similar differences. Conclusions: RLMs' reasoning can effectively assess complex case factors such as parental cooperation. Lower accuracy in assessing fathers' cooperation supports the argument of a stronger professional focus on mothers in CPS interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。