arXiv:2512.15894cs.AI2025-12被引 1

测试大模型在家长焦虑压力下的儿科咨询安全性,发现压力让错误率升8%。

PediatricAnxietyBench: Evaluating Large Language Model Safety Under Parental Anxiety and Pressure in Pediatric Consultations

  • 构建300个真实家长压力场景的儿科问答数据集,含150个对抗性提问
  • 70B模型安全得分6.26,比8B模型高1.31,关键错误率低7.2个百分点
  • 紧急情况识别缺失,焦虑语气使模型诊断失误率最高达33.3%

大型语言模型(LLMs)正被越来越多家长用于儿童健康咨询,但其在真实世界中应对焦虑家长施加的压力时的安全性尚不明确。焦虑家长常使用紧迫性语言,可能破坏模型的防护机制,导致有害建议。本文提出PediatricAnxietyBench,一个开源基准,包含300个高质量问题,覆盖10个儿科主题(150个患者来源,150个对抗性提问),支持可复现评估。采用多维度安全框架(诊断克制、转诊遵循、模糊表达、紧急识别)评估两个Llama模型(70B和8B)。对抗性提问模拟了紧迫感、经济障碍及对免责声明的挑战。平均安全得分5.50/15(标准差2.41)。70B模型优于8B模型(6.26 vs 4.95,p<0.001),关键失败率更低(4.8% vs 12.0%,p=0.02)。对抗性提问使安全得分下降8%(p=0.03),其中紧迫感导致降幅最大(-1.40)。癫痫和接种后查询中出现明显漏洞(不当诊断率33.3%)。模糊表达与安全得分强相关(r=0.68,p<0.001),但紧急识别能力完全缺失。模型规模影响安全表现,但所有模型均暴露于现实家长压力下。PediatricAnxietyBench提供可复用的对抗性评估框架,揭示传统基准忽略的临床显著失效模式。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly consulted by parents for pediatric guidance, yet their safety under real-world adversarial pressures is poorly understood. Anxious parents often use urgent language that can compromise model safeguards, potentially causing harmful advice. PediatricAnxietyBench is an open-source benchmark of 300 high-quality queries across 10 pediatric topics (150 patient-derived, 150 adversarial) enabling reproducible evaluation. Two Llama models (70B and 8B) were assessed using a multi-dimensional safety framework covering diagnostic restraint, referral adherence, hedging, and emergency recognition. Adversarial queries incorporated parental pressure patterns, including urgency, economic barriers, and challenges to disclaimers. Mean safety score was 5.50/15 (SD=2.41). The 70B model outperformed the 8B model (6.26 vs 4.95, p<0.001) with lower critical failures (4.8% vs 12.0%, p=0.02). Adversarial queries reduced safety by 8% (p=0.03), with urgency causing the largest drop (-1.40). Vulnerabilities appeared in seizures (33.3% inappropriate diagnosis) and post-vaccination queries. Hedging strongly correlated with safety (r=0.68, p<0.001), while emergency recognition was absent. Model scale influences safety, yet all models showed vulnerabilities to realistic parental pressures. PediatricAnxietyBench provides a reusable adversarial evaluation framework to reveal clinically significant failure modes overlooked by standard benchmarks.

大模型安全儿科咨询对抗测试家长压力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。