arXiv:2606.00027cs.CLcs.AI2026-06被引 1

构建多领域红队测试框架,评估医疗大模型在安全、鲁棒性与公平性上的真实表现。

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

论文配图:A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models
图 1 · 摘自论文原文
  • 设计跨9个领域、150+子类的临床场景红队测试框架
  • 高分模型仍存在致命安全漏洞,平均分0.791~0.984
  • 混合人机评估更可靠,可发现自动化遗漏的临床风险

大型语言模型(LLMs)正被广泛应用于医疗领域,但现有评测基准无法捕捉临床实践中常见的对抗性或伦理复杂情境。我们构建了一个多领域红队测试框架,评估了11个主流医疗大模型在690个基于临床的场景中的表现,覆盖9个领域及超过150个子类别。场景引入对抗性变换,响应通过七维度评分体系进行评估,结合大模型辅助打分与人工复核。结果显示性能差异显著,平均分介于0.791至0.984之间。关键发现:部分高性能模型在个别安全敏感场景中出现完全失效,表明均值准确率会掩盖临床风险。表现最优的模型(X-BAI、GPT-5、Claude Opus 4.1)得分超0.97且波动小,但不同领域间表现差异明显。涉及公平性的任务在引入人口学特征修改后,错误率提升10%-20%;人工评审识别出自动化评估未察觉的临床失误。研究证明,性能波动与最差情况下的失败比平均准确率更能反映临床可靠性,且融合自动化与临床专家监督的混合评估是可信安全评估的关键。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or ethically complex conditions common in clinical practice. We developed a multi-domain red teaming framework evaluating eleven contemporary LLMs across 690 clinically grounded scenarios spanning nine domains and over 150 subcategories. Scenarios incorporated adversarial transformations, and responses were assessed using a seven-dimension rubric with LLM-assisted scoring and human-in-the-loop validation. Results revealed substantial performance variance, with mean scores ranging from 0.791 to 0.984. Critically, several high-performing systems produced complete failures in individual safety-critical scenarios, demonstrating that aggregate accuracy masks clinically meaningful risk. The highest-performing systems (X-BAI, GPT-5, Claude Opus 4.1) achieved scores above 0.97 with low variance, while performance varied significantly across domains. Equity-related tasks showed 10-20% error amplification with demographic modifications, and human reviewers identified clinically relevant failures missed by automated evaluation. Our findings demonstrate that performance variance and worst-case failures provide more clinically meaningful reliability indicators than mean accuracy alone, and that hybrid evaluation approaches combining automation with clinician oversight are essential for credible safety assessment.

医疗大模型红队测试安全性评估公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。