临床AI在微小扰动和多语言环境下诊断准确率暴跌,暴露安全漏洞。
Adversarial Fragility and Language Vulnerability in Clinical AI: A Systematic Audit of Diagnostic Collapse Under Imperceptible Perturbations and Cross-Lingual Drift in Low-Resource Healthcare Settings
- 用FGM扰动测试模型,0.021幅度下准确率从89.3%跌至62.0%
- 英语、尼日利亚皮钦语、约鲁巴语变体中模型诊断一致率下降至50%
- 揭示低资源医疗场景下临床AI的脆弱性,警示需更鲁棒的模型设计
当前临床人工智能系统主要在干净、标准化的英文输入下评估,无法反映低资源医疗环境的现实。本研究首次系统性地审计了临床AI的两大安全缺陷:对抗图像脆弱性和跨语言诊断漂移。基于用于胸部X光检测的DenseNet121架构,在COVID-QU-Ex数据集(85,318张图像,含新冠、非新冠肺炎、正常)上微调后,发现当使用ε=0.021的快速梯度法(FGM)进行扰动时,诊断准确率从89.3%骤降至62.0%,且人类无法察觉该扰动。标准防御策略如高斯平滑与集成投票均未能恢复临床安全性。在另一语言脆弱性实验中,测试Llama3.1:8b与NatLAS模型在标准英语、尼日利亚皮钦语(Naija)及约鲁巴语化英语的20个新冠病例上的表现。两者均出现显著性能下降:Llama3.1:8b在皮钦语中准确率从80.0%降至65.0%;而针对非洲语境优化的NatLAS模型从85.0%跌至55.0%,诊断一致性仅剩50%。这些结果为尼日利亚基层卫生中心(PHC)部署条件下的临床AI设定了量化失效边界,亟需构建抗对抗攻击、语言包容性强的临床AI架构。
原文摘要 · Abstract (English)
Current clinical artificial intelligence (AI) systems are evaluated almost exclusively on clean, standardised, English-language inputs, conditions that do not reflect the realities of healthcare delivery in low-resource settings. This study presents the first systematic dual audit of two orthogonal safety vulnerabilities in clinical AI: adversarial image fragility and cross-lingual diagnostic drift. Using DenseNet121, the architecture underlying CheXNet, fine-tuned on the COVID-QU-Ex chest X-ray dataset (85,318 images; COVID-19, Non-COVID Pneumonia, Normal), we demonstrate that diagnostic accuracy collapses from 89.3% to 62.0% under a Fast Gradient Method (FGM) perturbation of epsilon=0.021, a magnitude imperceptible to the human eye. Standard defensive strategies including Gaussian smoothing and ensemble voting failed to restore clinical safety. In a parallel language fragility experiment, we tested Llama3.1:8b and NatLAS (N-ATLAS) on 20 COVID-19 clinical cases presented in Standard English, Nigerian Pidgin (Naija), and Yoruba-inflected English. Both models exhibited significant accuracy degradation: Llama3.1:8b dropped from 80.0% to 65.0% on Pidgin; NatLAS, an African-context model, collapsed from 85.0% to 55.0%, with diagnosis consistency falling to 50%. These findings establish a quantitative failure envelope for clinical AI under conditions representative of Primary Health Centre (PHC) deployment in Nigeria, and motivate urgent calls for adversarially hardened, linguistically inclusive clinical AI architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。