arXiv:2605.04039cs.CLcs.AI2026-05被引 4

医学大模型安全与准确遵循不同扩展规律,证据质量决定安全性。

Safety and accuracy follow different scaling laws in clinical large language models

论文配图:Safety and accuracy follow different scaling laws in clinical large language models
图 1 · 摘自论文原文
  • 构建SaFE-Scale框架,量化医疗大模型在不同部署条件下的安全表现。
  • 使用高质量证据使错误率从12.0%降至2.6%,但检索增强未改善高风险错误。
  • 适合关注临床应用安全性的医生、研发者及系统设计者阅读。

临床大模型常通过增大模型规模、上下文长度、检索复杂度或推理算力来扩展,隐含假设是更高准确率意味着更安全行为。但在医学领域,少数自信、高风险或违背证据的错误可能比平均表现更重要。本文提出SaFE-Scale框架,用于评估临床大模型在模型规模、证据质量、检索策略、上下文暴露和推理算力等维度下的安全性变化。为此,我们构建了RadSaFE-200——一个包含200道多选题的放射科安全评测基准,涵盖临床专家定义的清洁证据、冲突证据及选项级标签(高风险错误、不安全回答、证据矛盾)。对34个本地部署的LLM在六种部署条件下进行评估:闭卷提示(零样本)、清洁证据、冲突证据、标准RAG、代理式RAG和最大上下文提示。清洁证据使平均准确率从73.5%提升至94.1%,高风险错误从12.0%降至2.6%,证据矛盾从12.7%降至2.3%,危险自信心从8.0%降至1.6%。标准RAG与代理式RAG虽提升准确率并降低矛盾,但高风险错误与危险自信心仍居高不下。最大上下文提示增加延迟却未缩小安全差距,额外推理算力仅带来有限收益。最坏情况分析显示,关键临床错误集中于少数问题。因此,临床大模型的安全性并非扩展的自然结果,而是由证据质量、检索设计、上下文构建及集体失败行为共同决定的部署属性。

原文摘要 · Abstract (English)

Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation that higher accuracy implies safer behavior. This assumption is incomplete in medicine, where a few confident, high-risk, or evidence-contradicting errors can matter more than average benchmark performance. We introduce SaFE-Scale, a framework for measuring how clinical LLM safety changes across model scale, evidence quality, retrieval strategy, context exposure, and inference-time compute. To instantiate this framework, we introduce RadSaFE-200, a Radiology Safety-Focused Evaluation benchmark of 200 multiple-choice questions with clinician-defined clean evidence, conflict evidence, and option-level labels for high-risk error, unsafe answer, and evidence contradiction. We evaluated 34 locally deployed LLMs across six deployment conditions: closed-book prompting (zero-shot), clean evidence, conflict evidence, standard RAG, agentic RAG, and max-context prompting. Clean evidence produced the strongest improvement, increasing mean accuracy from 73.5% to 94.1%, while reducing high-risk error from 12.0% to 2.6%, contradiction from 12.7% to 2.3%, and dangerous overconfidence from 8.0% to 1.6%. Standard RAG and agentic RAG did not reproduce this safety profile: agentic RAG improved accuracy over standard RAG and reduced contradiction, but high-risk error and dangerous overconfidence remained elevated. Max-context prompting increased latency without closing the safety gap, and additional inference-time compute produced only limited gains. Worst-case analysis showed that clinically consequential errors concentrated in a small subset of questions. Clinical LLM safety is therefore not a passive consequence of scaling, but a deployment property shaped by evidence quality, retrieval design, context construction, and collective failure behavior.

医疗AI大模型安全RAG评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。