arXiv:2508.00923cs.LG2025-08被引 16

动态红队测试发现医疗大模型看似高分实则脆弱,暴露真实场景下安全短板。

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

  • 用自动变异测试案例的动态红队框架持续检测模型安全漏洞。
  • 94%高分答案在动态测试中失效,隐私、偏见、幻觉问题普遍超70%。
  • 适合医疗AI开发者和监管者,用于发现静态评测遗漏的真实风险。

大型语言模型(LLMs)被广泛用于回答健康问题并支持医疗流程,但其安全性评估仍依赖易过时的静态基准。本文提出动态、自动、系统化(DAS)红队审计框架,从鲁棒性、隐私、偏见/公平性、幻觉/事实错误四方面持续压力测试医疗领域模型。经认证临床医生验证,一组自主对抗代理能实时变异健康测试案例以揭示漏洞。对15个专有及开源模型应用DAS发现:尽管中位数MedQA准确率超80%,94%原正确答案在动态鲁棒性测试中失效;该脆弱性在开放式的HealthBench数据集上同样显著,顶尖模型失败率超70%,排名剧烈波动,表明静态高分可能反映浅层记忆。此外,86%场景出现隐私泄露,81%公平性测试中认知偏见改变推荐,主流模型幻觉率超74%。DAS将医疗大模型安全评估从静态清单转变为动态对抗审计,可规模化识别部署前潜在风险,适用于面向消费者的健康助手、医生工具等场景。代码已公开于https://github.com/JZPeterPan/DAS-Medical-Red-Teaming-Agents。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against. Here we introduce a Dynamic, Automatic, and Systematic (DAS) red-teaming audit framework that continuously stress-tests LLMs for health across four safety-critical axes: robustness, privacy, bias/fairness, and hallucination/factual inaccuracies. Validated against board-certified clinicians with high concordance, a suite of adversarial agents autonomously mutates health-related test cases to uncover vulnerabilities in real time. Applying DAS to 15 proprietary and open-source LLMs revealed a profound gap between high static benchmark performance and low dynamic reliability--the "Benchmarking Gap". Despite median MedQA accuracy exceeding 80\%, 94\% of previously correct answers failed under dynamic robustness testing. This brittleness generalized to the realistic, open-ended HealthBench dataset, where top-tier models exhibited failure rates exceeding 70\% and sharp shifts in model rankings across evaluations, suggesting that high scores on established static benchmarks may reflect superficial memorization. We observed similarly high failure rates across other domains: privacy leaks were elicited in 86\% of scenarios, cognitive-bias priming altered recommendations in 81\% of fairness tests, and hallucination rates exceeded 74\% in widely used models. By converting LLM safety evaluation for health from a static checklist into a living adversarial audit, DAS provides a scalable framework for surfacing latent risks before such systems are deployed in consumer-facing health assistants, clinician-facing tools, and broader healthcare workflows. Code is available at https://github.com/JZPeterPan/DAS-Medical-Red-Teaming-Agents.

医疗AI安全评估红队测试大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。