arXiv:2608.02617cs.CLcs.AI2026-08

医生偏爱的AI回复未必安全,偏好不能替代临床安全评估

Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

论文配图:Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
图 1 · 摘自论文原文
  • 用医生盲评对比判断模型偏好,结合多维度评分验证安全
  • 高偏好模型仍存在显著临床错误率,部分领域风险集中
  • 表面特征比安全指标更影响医生选择,需调整排名方式

我们利用MOOVE平台(由全球736名医生在28个国家贡献的26,804条双模型对比判断)评估临床医生对大语言模型输出的偏好是否可作为临床安全的有效信号。医生在[-2, +2]量表上打分,负值表示存在临床不安全或误导性内容。结果显示,模型在偏好排序中靠前,仍可能在‘无害性’和‘准确性’等关键维度出现大量实质性失败(得分≤-1),且这些失败在不同医学专科间分布不均,形成不可忽视的‘禁区’。分析表明,偏好受提示长度、拒答与升级行为影响,表面特征解释力略高于安全核心维度。约一半偏好投票未传递正向安全信号。我们提出一种结合偏好与评分反馈的临床调优排名方法,优于仅依赖布拉德利-特里强度的原始排序。研究呼吁将偏好与安全评价分离,直接报告安全失败率,并在临床决策场景中引入临床校准机制。

原文摘要 · Abstract (English)

We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings. Clinicians assign scores on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading content. Using 26{,}804 pairwise judgments across outputs from 13 LLMs, contributed by more than 736 clinicians across 28+ countries, we find that clinician preference is a poor proxy for safety-critical performance. Models ranking highly under pairwise preference can still exhibit substantial rates of clinically meaningful failures ($\leq -1$) on dimensions such as \emph{Harmlessness} and \emph{Accuracy}. These failures are unevenly distributed across specialties, creating domain-specific ``no-go zones'' not visible in aggregate rankings or single-number leaderboards. We further analyze contributing factors including prompt length, refusal and escalation behavior, and the relative contributions of safety-critical versus surface-level features. A substantial fraction of preference votes carry no positive safety signal, while feature decomposition shows that surface-level characteristics explain slightly more preference variation than safety-critical rubric differences. Finally, we introduce a clinically adjusted preference ranking combining pairwise preference with rubric-derived feedback, producing a more safety-aware ordering than raw Bradley--Terry strength alone. Our findings support evaluation practices that separate preference from safety, report safety-critical failure rates directly, and incorporate clinically grounded adjustments when ranking LLMs for clinical decision making.

大模型评估临床安全偏好学习医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。