arXiv:2607.18086cs.AI2026-07

测试提示词对临床大模型安全性的提升效果,发现收益依赖评分者且伴随帮助性损失。

Judge-dependent safety gains and model-specific helpfulness costs of evidence-sufficiency prompting in clinical LLMs

  • 用结构化提示降低模型在证据不足时的过度自信
  • 安全得分下降24.7个百分点,但不同评分模型结果差异大
  • 需结合人类医生评估,警惕评分者偏差和帮助性代价

背景:临床大模型在证据不全时是否过度自信,常由大模型评分器判断,但这种‘安全性提升’是真实行为改变还是评分器自身偏差尚不明确。本文以结构化证据充分性提示为测试案例,考察其能否减少不安全的过度自信回答,该效果是否依赖评分者,以及对帮助性的影响。方法:在三个公开数据集(Real-POCQi、HealthBench、MedRBench)上,四个模型(GPT-5.5、Claude Opus 4.8、Gemini 3.5 Flash、Grok 4.3)在1,200个问答对中使用标准提示与包装提示进行对比。主要终点为由主评分器(GPT-5.4-nano)评估的不安全过度自信率的配对下降;次要分析包括另一类评分器(Claude Sonnet 5)、正确性评分器、匹配对照组及三名临床医生盲评。结果:不安全过度自信从49.3%降至24.7%,配对下降24.7个百分点(95% CI 21.8–27.7;p<0.001),方向一致,跨模型与改写均稳健。但效果大小受评分者影响:Sonnet 5 同意方向但效应减半(+13.1点),存在单向分歧。盲评临床医生认为主评分器敏感度高(1.00)、特异度低(0.55),非校准评分。安全提升伴随模型特异性帮助性损失(正确诊断率从80.3%降至50.3%):GPT-5.5 几乎无成本,Gemini 损失达-58点。对照组显示行为变化为真实,非评分循环。结论:临床大模型安全性评估应报告方向性与相对性,锚定于人类评审,并联合评估帮助性,而非绝对校准值。本研究不意味着可直接部署临床使用。

原文摘要 · Abstract (English)

Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved. Using a structured evidence-sufficiency prompt as a test case, we asked whether it reduces unsafe overconfident answers, how far that effect depends on the scoring judge, and what it costs in helpfulness. Methods: In a retrospective public-data benchmark (Real-POCQi, HealthBench, MedRBench), four models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3) answered a fully paired common panel (1,200 cells) with a standard prompt and the wrapper. The pre-specified endpoint was the paired reduction in unsafe overconfidence scored by the primary judge (GPT-5.4-nano); secondary analyses added a different-family judge (Claude Sonnet 5), a correctness judge, matched scaffold controls, and a blinded three-clinician review. Results: Unsafe overconfidence fell from 49.3% to 24.7%, a paired reduction of 24.7 points (95% CI 21.8-27.7; p<0.001), robust in direction across models and paraphrases. Magnitude was judge-dependent: Sonnet agreed on direction but nearly halved the effect (+13.1 points), with one-directional disagreement. Blinded clinicians characterized the primary judge as a high-sensitivity (1.00), low-specificity (0.55) screen, not a calibrated rate. The gain carried a model-specific helpfulness cost (correct diagnosis 80.3% to 50.3%): near-free for GPT-5.5, near-total for Gemini (-58 points). Matched scaffold controls showed genuine behavior change, not judge circularity. Conclusions: LLM-judged clinical safety effects should be reported as directional and relative, anchored to human review and evaluated jointly with helpfulness, not as calibrated absolute rates. This does not establish clinical deployment readiness.

大模型安全临床应用评分偏差提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。