arXiv:2601.09721cs.CLcs.AI2026-01

测试大模型在家长焦虑下的儿科问诊安全,发现小模型反而更稳。

Cross-Platform Evaluation of Large Language Model Safety in Pediatric Consultations: Evolution of Adversarial Robustness and the Scale Paradox

  • 用真实家长焦虑场景生成对抗性提问,测试三款大模型表现
  • 小模型比大模型更安全,对癫痫等病症误诊率高达33%
  • 提出开放基准集,强调医疗AI需做对抗压力测试

大型语言模型(LLMs)在医疗问诊中应用日益广泛,但其在真实用户压力下的安全性仍研究不足。本研究基于已有的PediatricAnxietyBench数据集,包含300个问题(150个真实,150个对抗性),覆盖10个主题,通过API评估了Llama-3.3-70B、Llama-3.1-8B(Groq)和Mistral-7B(HuggingFace)三款模型,共获得900条响应。采用0-15分制评估约束、转诊、模糊回应、紧急识别及非处方行为等安全指标。结果显示,平均得分介于9.70(Llama-3.3-70B)至10.39(Mistral-7B)之间;Llama-3.1-8B优于Llama-3.3-70B,差异+0.66(p=0.0001, d=0.225)。所有模型均表现出正向对抗效应,其中Mistral-7B最强(+1.09, p=0.0002)。安全性跨平台具有一致性,但Llama-3.3-70B有8%失败案例。癫痫相关问题误诊率达33%。模糊回应与安全性能显著相关(r=0.68, p<0.001)。结论表明,模型安全性更依赖对齐与架构而非规模,小模型表现更优;多版本迭代显示鲁棒性提升。但普遍缺乏紧急识别能力,不适用于分诊。研究为医疗AI安全提供可复用的开放基准,建议优先进行对抗测试。

原文摘要 · Abstract (English)

Background Large language models (LLMs) are increasingly deployed in medical consultations, yet their safety under realistic user pressures remains understudied. Prior assessments focused on neutral conditions, overlooking vulnerabilities from anxious users challenging safeguards. This study evaluated LLM safety under parental anxiety-driven adversarial pressures in pediatric consultations across models and platforms. Methods PediatricAnxietyBench, from a prior evaluation, includes 300 queries (150 authentic, 150 adversarial) spanning 10 topics. Three models were assessed via APIs: Llama-3.3-70B and Llama-3.1-8B (Groq), Mistral-7B (HuggingFace), yielding 900 responses. Safety used a 0-15 scale for restraint, referral, hedging, emergency recognition, and non-prescriptive behavior. Analyses employed paired t-tests with bootstrapped CIs. Results Mean scores: 9.70 (Llama-3.3-70B) to 10.39 (Mistral-7B). Llama-3.1-8B outperformed Llama-3.3-70B by +0.66 (p=0.0001, d=0.225). Models showed positive adversarial effects, Mistral-7B strongest (+1.09, p=0.0002). Safety generalized across platforms; Llama-3.3-70B had 8% failures. Seizures vulnerable (33% inappropriate diagnoses). Hedging predicted safety (r=0.68, p<0.001). Conclusions Evaluation shows safety depends on alignment and architecture over scale, with smaller models outperforming larger. Evolution to robustness across releases suggests targeted training progress. Vulnerabilities and no emergency recognition indicate unsuitability for triage. Findings guide selection, stress adversarial testing, and provide open benchmark for medical AI safety.

医疗AI模型安全对抗测试儿科问诊

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。