arXiv:2605.28338cs.AI2026-05

通过医生审核的推理溯源,提升医疗大模型的安全与伦理对齐。

SafeMed-R1: Clinician-Audited Safety and Ethics Alignment for Medical Large Language Models

  • 构建可追溯的临床信任信号管道,关联医生评分与编辑历史。
  • 在临床基准上达79.6%准确率,对抗性测试中不安全输出降低3至5%。
  • 医生对比实验显示其用药安全性和临床实用性优于住院医师。

大型语言模型在执照考试中表现接近专家,但日常临床应用受限,因治理需可审计的推理、安全与伦理对齐,以及抗恶意滥用能力。本文提出SafeMed-R1,通过可追溯的临床信任信号(CTS)流程训练,将每条推理实例与医生评分及编辑历史关联,并经安全伦理监督与红队压力测试对齐。SafeMed-R1在临床基准上达到79.6%的宏平均准确率。在对抗性安全测试中,其累积风险最低,不安全输出相较基线减少约3至5%。在30个药物安全案例的配对专家研究中,SafeMed-R1在医学正确性上与PGY1、PGY2住院医师相当,且在用药安全性、指南一致性与临床实用性上得分更高。结果表明,医生审核的监督溯源结合领域定制化安全伦理对齐,可在无需推理时检索或引用验证的前提下,增强治理相关证据。

原文摘要 · Abstract (English)

Large language models(LLMs) increasingly match expert performance on licensing examinations, yet routine clinical use remains limited because governance requires auditable reasoning, safety and ethics alignment, and resilience to adversarial misuse. Here we present SafeMed-R1, trained with a traceable Clinical Trust Signals(CTS) pipeline that links each reasoning instance to clinician rubric scores and edit histories, and aligned through safety and ethics supervision and red team stress testing. SafeMed-R1 attains a macro-averaged accuracy of 79.6% across clinical benchmarks. Under adversarial safety testing, it shows the lowest aggregated risk and reduces unsafe outputs by about 3 to 5% relative to its baseline. In a paired expert study of 30 medication safety vignettes, SafeMed-R1 matches PGY1 and PGY2 residents on medical correctness and scores higher for medication safety, guideline consistency, and clinical usefulness. Collectively, these results suggest that clinician-audited supervision provenance, together with domain-tailored safety and ethics alignment, can strengthen governance-relevant evidence without relying on inference-time retrieval or citation grounding.

医疗LLM安全对齐临床可信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。