arXiv:2607.17270cs.CLcs.CY2026-07

小模型在本地部署时,医疗安全能力在英语到豪萨语间大幅下降。

Safety That Does Not Transfer: Cross-Lingual Clinical Correctness Drift in Deployable Medical Language Models

论文配图:Safety That Does Not Transfer: Cross-Lingual Clinical Correctness Drift in Deployable Medical Language Models
图 1 · 摘自论文原文
  • 用英语训练的医疗模型在豪萨语下临床正确性暴跌
  • 本地可部署模型在豪萨语中平均得分从1.57降至-0.03
  • 问题出在模型规模而非语言或疾病本身,适合低资源医疗场景研究者

大型语言模型的安全评估主要在英语环境下针对前沿模型进行,但真实低资源医疗场景中,小模型(4-9亿参数)在本地部署并以本地语言提问。本研究考察了英语中建立的临床安全性是否能迁移至豪萨语,以及失败原因是否来自语言、临床任务或部署模型类型。构建了三类高负担疾病的英豪双语问答对:疟疾、镰状细胞病、结核病,涵盖知识回忆、紧急分诊、诱导禁忌行为及传统疗法主张。评估六种模型:五种本地可部署系统(含两种医学微调),一种前沿模型。所有128条回复由两名流利豪萨语者独立盲评,依据尼日利亚国家治疗指南。本地模型在豪萨语中平均临床正确性从英语的1.57降至-0.03(满分2,-1为有害);前沿模型从2.00降至1.75,未产生有害回答。错误迁移在三种疾病中均一致。一致性分析显示临床正确性κ=0.70,有害性判断κ=0.22,需进一步分析。因前沿模型在豪萨语中表现良好,说明缺陷源于部署层级,而非语言或临床内容本身。

原文摘要 · Abstract (English)

Safety evaluation of large language models is conducted predominantly in English and predominantly on frontier systems. Neither condition describes how such models are encountered in low-resource health settings, where small quantised systems are run locally and queried in local languages. We ask whether clinical safety established in English transfers to Hausa, and whether any failure is attributable to the language, the clinical task, or the class of model that low-resource deployment admits. Matched English-Hausa question pairs were built for three conditions of high burden in northern Nigeria: malaria, sickle cell disease, and tuberculosis, probing knowledge recall, emergency triage, a leading question inviting a contraindicated action, and a traditional-remedy claim. Six models were evaluated: five locally deployable systems of 4-9 billion parameters, two medically fine-tuned, and one frontier system. All 128 responses were scored against Nigerian national treatment guidelines by two fluent Hausa speakers working independently and blind to one another. Among locally deployable models, mean clinical correctness fell from 1.57 in English to -0.03 in Hausa, on a scale where 2 denotes a correct answer and -1 an actively harmful one. The frontier model moved from 2.00 to 1.75 and produced no response judged harmful in either language. Drift was consistent across all three conditions. Inter-rater agreement was substantial for clinical correctness (kappa = 0.70); agreement on harm was initially poor (kappa = 0.22) and is examined in detail. Because a frontier model answers the same questions competently in Hausa, the deficit is a property neither of the language nor of the clinical material, but of the deployable tier.

医疗AI多语言模型安全低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。