用多模态融合无训练框架,帮孟加拉儿童筛查创伤,结果可解释且适配本地。
Explainable AI for Screening Abuse-Related Trauma in Bangladeshi Children: A Training-Free Multimodal Framework Evaluated on Noise-Aware Synthetic Data
- 四模态融合+临床权重,无需训练,支持单模态优先决策
- 在合成数据上达AUC 0.874,比仅用问卷提升显著
- 输出双语报告并对接保护机构,适合资源匮乏地区使用
孟加拉每10万人仅1.17名心理健康专业人员,全国仅有六名儿童精神科医生。目前缺乏本地化、文化适配的儿童虐待心理创伤早期筛查工具。本文提出ShishuRaksha AI,一种非诊断性决策支持框架,融合四种筛查模态:标准化量表(SDQ、CPSS)、孟加拉语叙事文本、房-树-人绘画特征及面部表情。该框架为无训练设计,采用跨模态注意力与临床加权,并设有单模态优先规则。所有风险评分通过扰动驱动的可解释性方法生成,以中英文报告形式呈现,并按《2013年儿童法》自动转介至国家儿童保护机构(OCC、DSS、NMHH)。因伦理限制,无法获取真实临床数据,故构建包含500例样本(116例阳性,占比23.2%)的噪声感知合成基准,设置四层人为噪声,基于文献的HTP先验知识。评估中排除面部通道,使用树集成代理模型,在五折分层交叉验证下,融合模型达到AUC 0.874(0.834–0.908),优于仅使用SDQ的基线模型(0.756,0.705–0.803)。通过消融、操作点、子群及校准分析验证有效性。研究明确指出局限性:仅依赖合成数据、无保留测试集、文本特征循环性以及城乡子群差距。本工作为低资源环境下可伦理部署的儿童保护筛查提供可行性探索与设计贡献。
原文摘要 · Abstract (English)
Bangladesh has an estimated 1.17 mental-health professionals per 100,000 population and only six child psychiatrists nationwide. No Bengali-language, culturally adapted tool exists for early screening of abuse-related psychological trauma in children. We present ShishuRaksha AI, a decision-support (not diagnostic) framework that fuses four screening modalities: validated questionnaires (SDQ, CPSS), Bengali narrative text, House-Tree-Person (HTP) drawing features, and facial affect. The fusion is training-free and clinically weighted, uses cross-modal attention, and includes a single-modality override rule. Every risk score is explained through clinically weighted, perturbation-based additive attribution and rendered as a bilingual (Bangla/English) report with referral routing to national child-protection services (OCC, DSS, NMHH) under the Children Act 2013. No clinical dataset of abused children can be collected ethically at this stage, so we introduce a noise-aware synthetic benchmark (500 cases, 116 positive [23.2%], four deliberate noise layers, literature-grounded HTP priors) and evaluate tree-ensemble surrogates of the fusion design (facial channel excluded) under 5-fold stratified cross-validation. The fused model reaches an AUC of 0.874 [0.834-0.908], against 0.756 [0.705-0.803] for an SDQ-only baseline, with ablation, operating-point, subgroup, and calibration analyses. We state all limitations openly, including synthetic-only data, no held-out set, text-feature circularity, and an urban-rural subgroup gap. This work is a feasibility study and a design contribution toward ethically deployable child-protection screening in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。